The situation
A knowledge-base assistant on Amazon Bedrock retrieves passages to ground its answers. The embeddings live in a vector store, and at launch the corpus was 40,000 ChunkingSplitting documents into retrievable pieces before embedding them – small enough to match precisely, big enough to still make sense.. Queries came back in a few milliseconds and recall was effectively perfect, because the store was doing an exact scan over every vector on every query. Nobody thought about the index because there wasn’t one worth naming.
Eighteen months later the corpus is 12 million chunks and growing, the embeddings are 1,024 Embedding dimensionHow many numbers each embedding vector holds – fewer means a smaller, cheaper, faster index and slightly blurrier matching., and the same exact scan now takes over a second per query. Retrieval has become the slowest part of the request. The obvious lever, a larger instance, gives a little headroom and then the curve catches up again, because exact search cost grows with the corpus and no amount of hardware changes that shape.
The store on offer, whether that’s Amazon OpenSearch Service with its k-NN plug-in or Aurora PostgreSQL with pgvector, supports ANNIndex structures (HNSW graphs, IVF partitions) that answer the k-nearest-neighbours question fast by giving up guaranteed exactness – recall becomes a tunable knob rather than a certainty. indexes. Switching to one will make queries fast again. The question underneath is which index, and what it gives up in exchange: a percent or two of recall, a chunk of memory, a longer build, or all three. Get it wrong and the assistant answers slowly, answers from the wrong passages, or runs a bill nobody signed off.
What actually matters
Vector search is a three-way tension, and every index choice is a point inside it. The three corners are recall (how often the approximate search returns the same neighbours an exact search would), latency (how fast a query comes back), and memory or cost (how much RAM and storage the index needs to hold). You cannot max all three at once. Exact search sits at the perfect-recall corner, where latency grows with the corpus; the approximate indexes cut latency by giving up a controllable slice of recall, and they differ mainly in how much memory they demand to do it.
Corpus size decides whether you even have a problem. At tens of thousands of vectors, an exact scan is fine and an index is premature; the scan is fast and its recall is a guaranteed 100%. The exact scan’s cost grows with the number of vectors, so somewhere between hundreds of thousands and a few million, depending on dimension and latency budget, the scan crosses from “instant” to “the bottleneck”. Approximate indexes exist to break that link, so their query cost grows far more slowly than the corpus does. The decision to index is really a decision about where you are on that curve.
Recall is a dial rather than a fixed property of the index. Every approximate index has parameters that trade recall against speed and memory, and the same index can be tuned to 99% recall or 90% recall on the same data. That means “which index” and “how is it tuned” are one question, not two. An HNSW index with a low search parameter can return worse results than a well-tuned IVF index, and vice versa. Quoting an index’s recall without quoting its parameters is meaningless.
Build cost and update cost are separate from query cost, and easy to forget until they bite. A graph index that answers queries beautifully can take hours to build and rebuild, and some index types need a training pass over a sample of the data before they can be populated at all. If the corpus changes constantly, the cost of keeping the index current can dominate the cost of querying it. A store that indexes 12 million vectors nightly has a very different profile from one that ingests a steady trickle.
Memory is often the real budget, and it is worth being precise about where it goes. Almost all of it is the vectors themselves. OpenSearch estimates a faiss HNSW index at roughly 1.1 * (4 * dimension + 8 * m) bytes per vector, so at 1,024 dimensions and m of 16 the raw floats are 4,096 bytes and the graph links only 128. A faiss IVF index over the same vectors comes out at roughly 1.1 * (4 * dimension) bytes per vector plus the centroids. Swapping a graph for clusters therefore saves a few percent of RAM, not a factor of anything. The only structural way down is to store fewer bits per vector, which is what quantisation does, and that shows up as lower recall.
None of this is answerable from a spec sheet. Recall depends on your embedding distribution, your query distribution, and your parameters, all of which are specific to your data. The only trustworthy numbers come from building a small ground-truth set (exact-search results for a sample of real queries) and measuring approximate recall and latency against it. Every recommendation below is a starting point to measure from, not a setting to trust blind.
What we’ll filter on
- Corpus scale, are we at tens of thousands of vectors where exact search is fine, or millions where an approximate index becomes necessary?
- Recall target, how close to exact-search results does retrieval need to be, and how much drop is tolerable?
- Query latency budget, what per-query time does the request path allow for the search step?
- Memory and cost ceiling, how much RAM is the index allowed to consume, and does that force compression?
- Build and update profile, is the corpus static, batch-rebuilt, or continuously changing, and what index-maintenance cost does that imply?
The landscape
Exact / brute-force (flat). No approximation: the query is compared against every vector and the true nearest neighbours come back. Recall is 100% by definition, there are no parameters to tune, and there’s nothing to build beyond storing the vectors. In pgvector this is simply a query with no ANN index present; in OpenSearch it’s exact k-NNThe retrieval question itself: given a query vector, return the k closest vectors under the index’s distance metric – answered exactly by comparing against everything, or quickly by an ANN index. scoring. The cost is linear in the corpus, so query time grows with the number of vectors. Perfect for small or slowly-searched corpora, and the ground truth you measure every other index against, but it stops scaling exactly when you need it to.
HNSW (Hierarchical Navigable Small World). A graph index: vectors become nodes connected to their near neighbours across several layers, and a query greedily walks the graph from an entry point toward the closest matches. Queries are very fast and recall is high, which is why it is the index the Bedrock Knowledge Bases setup documentation lands on for both back ends: faiss HNSW for an OpenSearch Serverless collection, and a USING hnsw pgvector index on Aurora. Memory and build time are what you give up. The graph plus the vectors generally live in RAM, and building the graph is slower and heavier than clustering-based alternatives. Three parameters do the tuning: m, the number of neighbour links per node (higher means better recall and more memory); ef_construction, the size of the candidate list while building (higher means a better graph and a slower build); and ef_search, the size of the candidate list at query time (higher means better recall and slower queries). The first two are fixed at build time; ef_search you can turn per query.
IVF / IVFFlat (inverted file). A clustering index: a training pass runs k-means over a sample to partition the space into nlist cells, each vector is assigned to its nearest cell centroid, and a query only scans the vectors in the few cells closest to it. Build is much faster than HNSW because there’s no graph to construct, just centroids and cell assignments. Memory is only slightly lower, since the full-precision vectors are still stored; the saving is the graph links alone. Recall is typically a touch lower for the same effort. The tuning dial is nprobe (spelled nprobes in OpenSearch, ivfflat.probes in pgvector), the number of cells a query scans: 1 is fast and low-recall, raising it scans more cells for better recall and more work per query, and at nprobe equal to nlist you’re back to an exact scan. The catch is training. IVF needs a pass over a representative sample before it can be populated, which on OpenSearch means building a model through the Train API first, and if the data distribution shifts a long way from that sample the cells stop being balanced and recall drifts.
Quantisation, layered on top. Not an index on its own but a way of storing each vector in fewer bits, pairable with either method. Product quantisation (PQ) splits a vector into sub-vectors and replaces each with the nearest entry in a learned codebook, so a 1,024-dimension float vector shrinks to a short code; like IVF, it has to be trained first. Binary quantisation, which faiss on OpenSearch has carried since 2.17, stores 1, 2 or 4 bits per dimension for 32x, 16x or 8x compression, and needs no separate training pass. Recall drops either way, because the stored vectors are now approximations. The standard remedy is a two-stage read: the compressed index shortlists candidates, then a RerankingA second pass that re-scores a wide set of retrieved candidates and keeps only the few most relevant, so the expensive model reads less. pass rescores the shortlist against full-precision vectors.
Where the stores sit. On an OpenSearch Service domain the k-NN plug-in offers HNSW on the faiss and lucene engines and IVF on faiss only, with the quantisation encoders also on faiss. The older nmslib engine is deprecated and blocked for new indexes from OpenSearch 3.0, so leave it out of the shortlist. OpenSearch Serverless narrows the choice further: its vector collections support HNSW with faiss and not IVF at all, and new-generation collections index at 32x compression by default. pgvector offers both HNSW and IVFFlat on a Postgres column, with m and ef_construction at build and hnsw.ef_search per session for HNSW, lists at build and ivfflat.probes per session for IVFFlat. The vocabulary differs; the three corners don’t.
Evaluation
Side by side
| Index | Recall | Query latency | Memory | Build cost | Key tuning dial | Best when |
|---|---|---|---|---|---|---|
| Exact / flat | ✓ (100%) | ✗ (grows with corpus) | Full vectors, no index | ✓ (none) | none | Small corpus, or ground truth |
| HNSW | ✓ (high) | ✓ (very fast) | ✗ (full vectors plus graph) | ✗ (slow, heavy) | ef_search, m |
Fast high-recall at scale, memory available |
| IVF / IVFFlat | Slightly lower | ✓ (fast) | ✗ (full vectors, a few % under HNSW) | ✓ (fast, needs training) | nprobe |
Rebuild window is the constraint, small recall loss acceptable |
| Quantised (binary or PQ) | ✗ (lowest before re-ranking) | ✓ (fast) | ✓ (8x to 32x smaller) | PQ trains, binary doesn’t | compression level, re-rank depth | Memory is the binding constraint; very large corpora |
The solution
For the assistant at 12 million vectors, the exact scan has to go; the only question is what replaces it, and the honest answer starts with measurement, not a default. Build a ground-truth set first: take a few hundred real queries, run them through the existing exact search, and record the true top-k for each. That’s the yardstick. Every candidate index gets scored on recall against that set and on p95 query latency, on the real corpus, before anything ships.
HNSW is the strong default when the recall target is high and the memory budget can absorb it. Start with moderate parameters, an m of 16 and an ef_construction in the low hundreds (AWS suggests 256 for Aurora on pgvector 0.6.0 and later, which builds indexes in parallel), then sweep ef_search at query time and watch recall and latency move together. Raising it climbs toward exact-search recall and adds milliseconds, and there is usually a knee where recall flattens and further increases only add latency. Set ef_search there. Check memory before committing: at 12 million vectors of 1,024 dimensions the formula puts the index near 55 GB, so size the instance to hold it, because an index that spills out of RAM loses the latency it was built for. m and ef_construction are baked in at build time, so a sweep that calls for a denser graph means a rebuild.
IVF is the pick when HNSW’s build profile is the problem, and expect no meaningful memory relief from it. If the corpus is rebuilt on a schedule and graph construction is stretching the window, IVFFlat clusters far faster over the same vectors. pgvector’s own guidance sets lists at the number of rows divided by 1,000 below a million rows and the square root of the row count above it, which puts 12 million chunks near 3,500; probes starts at the square root of lists. Train on a representative sample, then tune nprobe the way you tuned ef_search, low for speed and higher for recall, measured against the ground truth. Watch for distribution drift: if the corpus grows away from the training sample the cells go lopsided, recall sags, and the fix is a retrain rather than a reindex.
Quantisation is the only lever that moves memory by an order of magnitude, so it enters when RAM is the binding constraint rather than a line item. On faiss, binary quantisation is the simpler of the two: 1, 2 or 4 bits per dimension for 32x, 16x or 8x compression, applied during indexing with no training pass to maintain. PQ gets you to a similar place with a trained codebook, which is more machinery for the same shape of result. Either way raw recall drops, and the recovery is a re-ranking pass: the compressed index shortlists a few hundred candidates, and those get rescored against full-precision vectors. Use it when the numbers force the issue; below that, the recall you give up is a bad trade.
Whichever index lands, the dials are conceptually the same on OpenSearch k-NN and on pgvector, and the settings are not portable between corpora. An ef_search or nprobe that hit 98% recall on someone else’s data is a guess on yours until you’ve measured it. This is one component in a larger retrieval system, and the surrounding choices about where the vector store lives shape which of these indexes is even on the table.
Worked example
Take the assistant’s numbers: 12 million chunks, 1,024-dimension embeddings, a per-query latency budget of about 50 ms for the search step, and a recall target of 95% against exact search. Exact scan currently runs over a second, so it’s out.
First, the ground truth. Sample 300 production queries, run exact k-NN for each, store the true top-10. That set never changes and every measurement below scores against it.
HNSW attempt. Build with m = 16, ef_construction = 200. The estimate of 1.1 * (4 * 1024 + 8 * 16) bytes a vector puts 12 million of them near 55 GB, so the instance is sized to keep the index resident. Sweep ef_search: at 40 the sample shows roughly 93% recall at around 8 ms; at 100, roughly 97% at around 18 ms; at 200, 98% at around 35 ms. The knee is near ef_search = 100, comfortably inside the latency budget and past the 95% target. If memory at this size is affordable, HNSW ships here and there’s no reason to give up the recall.
IVF alternative, run in parallel because the nightly rebuild window is tight. Train on a 1-million-vector sample, nlist = 4,096, which is close to the square root of the corpus. Sweep nprobe: at 16, about 91% recall at around 6 ms; at 64, about 96% at around 14 ms; at 128, about 97% at around 24 ms. nprobe = 64 clears the target inside budget and the build finishes far sooner than the HNSW graph. Memory lands within a few percent of HNSW, so this is a rebuild-window fix rather than a bill fix, and it adds a training step to maintain.
Memory-pressed variant. If holding 12 million full-precision vectors is the line that breaks the budget, faiss quantisation shrinks the resident footprint by up to 32x, at a raw recall somewhere in the 80s. Add a re-rank: shortlist 200 candidates from the compressed index, rescore them against full-precision vectors held outside RAM, and measured recall on the top-10 climbs back over 95%. More moving parts, far less memory, target still met. Run all three against the one ground-truth set and the pick stops being an opinion and becomes a number you can defend.
What’s worth remembering
- Recall is a tuned dial. Vector search trades recall, latency and memory; a recall figure means nothing without its parameters.
- Scale decides whether to index. Exact search suits tens of thousands of vectors; its cost grows with the corpus, so millions need an index.
- HNSW is the high-recall default. It builds slowest;
mandef_constructionare fixed at build time,ef_searchtunes recall per query. - IVF trades recall for build speed.
nprobeis the dial; it needs a training pass, and OpenSearch Serverless does not support it. - Quantisation is the memory lever. Binary quantisation compresses 8x to 32x; recall drops, and re-ranking against full-precision vectors recovers it.
- Measure on your own data. Build a ground-truth set from real queries and score recall and latency before committing to an index.