Vector stores — study guide
The concept's fragments, read in order.
Search by meaning, not by match
An ordinary keyword search matches text to text: ask for dog and you get the rows that contain the letters dog, and nothing else. A document about a puppy or a terrier sits right there and never surfaces, because the words never line up. Matching is exact and literal, and it is blind to everything that means the same thing in different words.
A vector store answers a different question. It keeps embeddings — the lists of numbers that stand for the meaning of items — and, given a query turned into its own vector, it returns the stored vectors nearest to it. Nearness stands in for similarity of meaning, so a search for dog can pull back the puppy and the terrier even though not one character matches. You are asking what is most similar, not what is spelled the same.
Everything here is how that works and what it costs: how an item becomes a point in a space you can measure distances in, how nearness is turned into a single number, how the closest points are found, and the indexes that keep the search fast when the stored points number in the millions.
A point in a space of meaning
An embedding is a list of numbers — a vector — that places an item at a point in a space, arranged so that items meaning similar things land near each other. That is the whole property a vector store leans on: the coordinates already carry the meaning, and being close in the space means being close in meaning like a map where every place is set down so that things alike sit close together, and searching means asking what lies near a spot.
The space has many dimensions, far more than the two or three you can picture, but the intuition survives the squeeze. Each item is a point; each query, once it is turned into a vector by the same method, is a point too. Searching is then a geometric act rather than a textual one: drop the query somewhere in the space and look at what surrounds it.
So a vector store never reads the item to compare it. It compares coordinates, and the coordinates are the only thing it stores and searches over. Get the points placed well and the rest of the machinery — measuring distance, finding the closest — is just geometry.
Measuring how close two vectors are
Nearness has to become a single number before a machine can rank one match above another. There are a few standard ways to produce that number, and a vector store commits to one of them when it is set up.
Cosine similarity measures the angle between two vectors: point in the same direction and the score is high, no matter how long either vector is, so it compares orientation rather than magnitude. The dot product (or inner product) multiplies the vectors together component by component and sums the result, folding both direction and length into one figure. Euclidean distance, also called L2 distance, is the plain straight-line gap between the two points — the everyday sense of how far apart they are. The first two grow as vectors become more alike; the distance shrinks, so smaller is closer.
The choice of measure decides what counts as similar, which is why a store fixes it up front rather than per query. Systems that expose these directly show all three side by side: pgvector, for instance, offers a cosine operator, an inner-product operator, and an L2-distance operator, one for each notion of close.
Find the k closest
The core operation of a vector store has a name: k-nearest-neighbor search. Give it a query vector and a number k, and it returns the k stored vectors closest to the query under whatever similarity measure the store uses. Not the single best match but the handful of best matches, ranked, because the runner-up often matters as much as the winner like finding the handful of houses closest to a point on a map, rather than looking up one exact address.
Done exactly, this is brute force. There is no shortcut that reads the query and knows which stored points are near it, so the store computes the distance from the query to every stored vector, then keeps the k smallest. The work is linear in the size of the collection: twice the vectors, twice the comparisons.
For a few thousand vectors that is nothing. For millions of them, every single query sweeping the entire collection is the wall you hit — and it is the reason the store reaches for an index instead of measuring against everything.
Almost the nearest, far faster
Comparing a query against every stored vector is honest and, past a certain scale, hopeless. The escape is to stop insisting on the exact answer. An approximate nearest neighbor index — an ANN index — returns neighbors that are almost always the true closest ones, while looking at only a small slice of the collection instead of all of it like a shortcut that returns almost exactly the right neighbors far faster than checking every point one by one.
Two families are common. HNSW builds a layered graph you can hop across, descending toward the query's neighborhood in a few jumps rather than scanning. IVF groups the vectors into clusters ahead of time, then at query time searches only the clusters nearest the query and skips the rest. Both reach a small candidate set fast and measure the query against just those.
The price is recall: the fraction of the true nearest neighbors the search actually finds. An ANN index may occasionally miss one, and both families let you tune how hard they look, buying back accuracy with more work or spending accuracy for more speed. That trade is the entire point — a slightly imperfect answer now beats a perfect answer that arrives too late.
Similarity plus a where-clause
A vector rarely travels alone. Alongside each one a vector store keeps metadata — ordinary key-value fields such as the owner, the source document, a timestamp, a category. The vector answers what is similar; the metadata says everything else about the item that a plain field can hold.
Keeping both together lets one query ask two questions at once. A hybrid query combines the similarity search with an ordinary filter, so you can ask for the most similar items that also satisfy a condition: the nearest matches among only this user's documents, or only among items from the last week. The filter narrows the field to the rows allowed to compete, and the similarity search ranks what survives.
Without this the store would be a blunt instrument, always searching the whole collection and unable to respect who owns what or when it arrived. The metadata is what turns similarity search into something you can point at a specific, permitted slice of the data.
When a table is not enough
A regular database is built to answer what exactly matches: the row whose id is 7, the orders placed today, the users named Lee. Its indexes are sorted structures that jump straight to values you can compare and order. Ask it instead for the rows most similar in meaning to a phrase, and it has no notion of what that means — there is no column to sort meaning by.
A vector store is built for exactly that missing question. It ranks stored items by nearness in the embedding space and hands back the most similar, which is a kind of answer a table's ordinary indexes cannot produce. The two systems are good at different questions: precise lookups on one side, similarity ranking on the other.
The line between them is not fixed, though. Some relational databases now add a vector column type and vector indexes so they can do both — pgvector gives PostgreSQL a vector type and ANN index types such as HNSW and IVFFlat, letting one database run a similarity search beside its ordinary queries. You do not always have to choose a separate store; sometimes the table learns the new trick.
Giving an agent a memory to search
A large collection of documents is useless to a system that can only match keywords through it. The questions people actually ask rarely share their exact words with the answer buried in the text, so literal matching walks right past the passage that would help. Similarity search is how you reach that passage anyway — by meaning, not by spelling.
This is what makes a vector store load-bearing for an agent. An agent pipeline turns its own documents into embeddings, stores them, and, faced with a task, searches the store to pull the handful of chunks most relevant to what it is doing before it acts on them. The vector store is the searchable memory that finds the right pieces out of far too many to read.
A few pieces make that search trustworthy and fast: nearness stands in for relevance, k-nearest-neighbor returns the best few, and an ANN index keeps the lookup quick across millions of chunks. Get those right and an agent can find what it needs in its own memory the moment it needs it.