Exam Room · Advanced Generative AI Developer

How Many Chunks to Retrieve: Tuning Top-K

· 28 min read

Generative AI Development · part of The Exam Room

The situation

A team has a working retrieval-augmented-generation assistant over a knowledge base of a few thousand support articles and product docs. The documents are chunked, embedded, and stored in a vector index; at query time the retriever pulls the nearest ChunkingSplitting documents into retrievable pieces before embedding them – small enough to match precisely, big enough to still make sense. by embedding similarity, pastes them into the prompt as context, and a Claude model on Amazon Bedrock answers from them. The pipeline was stood up quickly, and the retriever returns whatever the starter template set: top-k of 3.

Two complaints have arrived from different directions. Support engineers say the assistant sometimes returns “I can’t find that” for an answer plainly written in an article they can point to. Other times it returns a specific detail that appears in no document at all. Separately, finance has noticed the Bedrock input-token bill climbing after someone bumped k to 20 to fix the first complaint, and p95 latency roughly doubled. The “cannot find it” answers did not go away, and a few new wrong answers appeared.

The knob in the middle of both complaints is the same one: how many chunks the retriever hands the model on each call. Nobody has measured what the right number is; it has been guessed twice and guessed wrong twice.

What actually matters

Top-k is a recall-versus-precision-and-cost trade-off, and both ends of the range fail in their own way. Start with what a higher k does for recall. Retrieval by embedding similarity is imperfect: the chunk that literally contains the answer is not always the single Nearest-neighbour searchFinding the vectors closest to a query vector; at scale it’s approximated, trading a little accuracy for a lot of speed., because wording differs, the question is phrased unlike the source, or several chunks look similar. Raising k widens the net, so the answer-bearing passage is more likely to be somewhere in the set. If recall fails, nothing downstream can recover: the model cannot cite a passage it never received, so the response is either a refusal or an answer no retrieved passage supports. Many RAG answers written off as hallucination are retrieval misses.

So more chunks always helps recall. The reason not to simply set k to 50 is that every extra chunk has three effects at once. It adds input tokens on every call, and in a RAG prompt the retrieved context is usually far larger than the question and the answer put together. It adds latency, because a longer prompt takes longer to process. And it can lower answer quality, because a bigger context is not a neutral container. Padding the prompt with lower-relevance chunks dilutes the signal, so the one good passage sits among distractors and the response may be drawn from a plausible-looking but wrong chunk. Published work on long-context models, Liu and colleagues in “Lost in the Middle”, found accuracy highest when the relevant passage sits near the start or the end of the context and lowest when it sits in the middle, even for models built for long contexts. A relevant chunk at position 12 of 20 can go unused although it was retrieved.

That gives the shape of the curve. As k rises from very low, answer quality climbs steeply, because you are rescuing answers that were being missed for lack of the right passage. It plateaus once the answer-bearing chunk is reliably in the set. Then, as k keeps rising, quality sags, because extra chunks add distractors rather than coverage, while cost and latency keep climbing. The best k sits at the knee: high enough to clear the recall problem, low enough to stay out of the dilution zone. Where that knee falls is specific to the corpus and the chunking, so measure it rather than inherit it from a template.

Chunk size is coupled to k, and it moves the knee more than anything else does. Bedrock Knowledge Bases splits content into chunks of roughly 300 tokens by default, honouring sentence boundaries, and you can configure fixed-size, hierarchical or semantic chunking instead. Small chunks are precise but each holds little, so an answer spanning a couple of paragraphs may need several chunks retrieved together, which pushes the right k higher. Large chunks carry more context each, so fewer of them cover an answer and k can be lower, but each one uses more tokens and drags in off-topic text around the relevant sentence. A change to chunk size moves the whole curve, so the two are tuned together against the same eval set.

There is a way to reach high recall without sending a big k to the generator. Retrieve a wide net, then re-rank and keep only the best few for the model. The vector search returns, say, the top 25 candidates; a reranker model scores each chunk against the query and reorders the list by that score; only the top 4 or 5 go into the prompt. Recall comes from the wide first pass, precision from the reranker, and the model receives a short, high-signal context. Bedrock Knowledge Bases supports this on both the Retrieve and the RetrieveAndGenerate call, so the wide-then-narrow shape is configuration rather than custom code.

What we’ll filter on

  1. Recall at k, is the answer-bearing chunk actually in the retrieved set often enough on your own questions?
  2. Answer quality, do the model’s answers get better or worse as k changes, judged on an eval set rather than by feel?
  3. Cost per call, how many input tokens does this k add to every query, and is the quality gain worth that?
  4. Latency, what does the added context do to p95 response time?
  5. Chunk-size coupling, is the right k being set for the chunk size actually in use, or inherited from a different one?
  6. Two-stage option, would a wide retrieve plus a reranker reach the same recall with a shorter prompt?

The landscape

Very low k (1 to 2). Cheapest and fastest, and fine when chunks are large and self-contained or the corpus is tiny and each answer lives in one obvious place. The failure mode is recall: any question whose answer is not the single nearest neighbour gets a miss, and misses turn into refusals or unsupported answers. On a general knowledge base this is usually too tight.

Moderate k (3 to 8). The working range for most RAG systems with sensibly sized chunks. Bedrock Knowledge Bases sits in this band out of the box, returning up to five source chunks unless you set numberOfResults, which accepts 1 to 100. There is enough coverage that the answer passage is usually present, without so much padding that dilution and cost dominate. The exact figure inside the band is worth measuring, because 4 and 8 can differ noticeably in both quality and bill.

High k (10 to 20 plus). Maximises recall and is defensible when chunks are small so an answer needs several to be complete, or when a downstream reranker will trim the set before it reaches the model. Passed raw to the generator, it exposes the position effect above and runs up the largest token bill, and past the knee it lowers answer quality rather than raising it. Treat a high k as a way to reach recall, then cut the set back before generation.

Wide retrieve, then rerank. Retrieve many candidates, score them with a reranker, keep the best few for the model. This separates recall from prompt length: the first pass is wide, and the model receives only a short high-signal context. Bedrock charges reranking per query rather than per chunk, and one query covers up to 100 document chunks, so widening the first pass from 25 to 100 does not change the reranking charge. It adds one model call to the path. The strongest general answer when a single fixed k cannot satisfy both recall and precision at once.

Dynamic / threshold-based k. Rather than a fixed count, keep every chunk above a similarity score, so easy queries with one strong match return few and broad queries return more. Retrieval depth then follows the question, but a raw similarity threshold is hard to set well and varies by embedding model, so it usually needs a reranker’s calibrated relevance scores to be dependable.

Evaluation

Side by side

Approach Recall Answer precision to model Token cost Latency Best when
Very low k (1-2) ✗ ✓ Lowest Lowest Large self-contained chunks, tiny corpus
Moderate k (3-8) ✓ ✓ Low-medium Low Most RAG with sensible chunk sizes
High k (10-20+) ✓ ✗ High High Small chunks, or a reranker trims after
Wide retrieve + rerank ✓ ✓ Medium Medium Recall and precision both needed at once
Threshold-based k ✓ (varies) ✓ (varies) Varies Varies Query difficulty varies widely, scores calibrated

Reading the table against the team’s problem: the starting k of 3 was risking recall on a general knowledge base, and the jump to 20 swapped that miss for dilution, latency, and cost without fixing it, because the answer-bearing chunk was now present but buried. The row that resolves both is the wide-retrieve-then-rerank one, or a measured moderate k if a reranker is not available.

The solution

Start by measuring recall, because it is the failure most often mistaken for hallucination and the one you can quantify cleanly. Build a small eval set of real questions paired with the passage that answers each. Run retrieval at several values of k and record how often the answer passage appears anywhere in the returned set. That curve tells you the smallest k that clears the recall problem for your corpus. If recall is still poor at high k, more chunks will not fix it: the retriever is returning the wrong things rather than too few of them, and the repair is the embedding model, the chunking, or a reranker.

With recall understood, tune for the knee rather than the ceiling. Above the k where recall plateaus, extra chunks stop adding coverage and start adding distractors, so answer quality flattens and then declines while cost and latency keep rising. Judge answer quality with an evaluation harness (an LLM-as-a-judgeUsing a second model, prompted with a rubric, to score another model’s output when there’s no exact answer to diff against. grade or exact-match against expected answers) across a sweep of k values, and pick the lowest k that sits on the quality plateau. That k has the recall without the dilution. It is a per-corpus number; a value copied from another system’s blog post is a guess.

Tune chunk size and k together, because moving one moves the other’s best value. Shrink chunks for precision and expect to raise k so a multi-paragraph answer is still covered; enlarge chunks and expect to lower k while each chunk uses more tokens. Re-running the recall-and-quality sweep after any chunking change keeps the two aligned. Hierarchical chunking needs one extra caution: numberOfResults counts child chunks, and the knowledge base replaces children with their shared parent chunk before returning, so a response can hold fewer results than you asked for and far more text than the count suggests.

When a single fixed k cannot give both recall and precision, reach for the two-stage pattern. On Bedrock, numberOfResults sets the width of the first pass and numberOfRerankedResults sets how many survive it, both in the range 1 to 100, on either the Retrieve or the RetrieveAndGenerate call. Three constraints are worth knowing first. Reranking covers textual data only. The reranker models are Region-limited: Amazon Rerank 1.0 (amazon.rerank-v1:0) runs in ap-northeast-1, ca-central-1, eu-central-1 and us-west-2, while Cohere Rerank 3.5 (cohere.rerank-v3-5:0) adds us-east-1 and is the only one of the two available there. And the Rerank quota is 10 requests per second, which throttles a busy assistant sooner than the 20 Retrieve requests per second do. The CloudWatch SearchUnits metric counts reranker queries, so the added charge is measurable from the first day. If you run the second stage yourself, the standalone Rerank API takes one query and up to 1,000 source documents per call.

Worked example

The team builds an eval set of 120 real support questions, each tagged with the article passage that answers it, and runs a sweep. At k of 2, recall is 71%: nearly a third of questions never receive their answer chunk, which lines up with the “it says it cannot find it” complaints. Recall climbs to 88% at k of 4, 95% at k of 8, and 97% at k of 15, flattening after that. Recall is essentially solved by k of 8, and k of 2 was the first mistake.

Now the quality sweep, graded by an LLM judge against reference answers. Answer quality rises with recall up to k of 8, then dips once the set goes to the model unfiltered. At k of 15 several answers are drawn from a plausible but wrong chunk, and a couple that were correct at k of 8 regress because the relevant passage now sits in the middle of a long context. Input tokens per call at k of 15 are nearly four times those at k of 4, and p95 latency is up by half. That is the k of 20 experiment, quantified: recall was fine, and dilution, latency, and cost all got worse together.

Answer quality and cost against top-k As top-k rises, recall and answer quality climb to a plateau near k of 8, then answer quality sags while cost keeps rising, marking a best k at the knee. low high value top-k (chunks retrieved) 1 4 8 15 25 best k (the knee) answer quality cost and latency

The two curves settle it. Set k at the knee, around 8, if the context goes straight to the model. Better, retrieve a wide net of 25, add the Knowledge Bases reranker, and pass the top 4 or 5: recall comes from the wide pass, the reranker ranks the genuinely relevant chunks first, and the prompt lands near the k of 4 token count rather than the k of 15 one. The reranking charge is one query either way, whether the first pass returned 25 chunks or 100. The “cannot find it” answers stop because recall is solved, and the wrong answers drop because the prompt no longer carries fifteen distractors.

What’s worth remembering

  1. Top-k trades recall for dilution. Too low misses the answer chunk; too high adds tokens, latency and distractors.
  2. Default k is five. Knowledge Bases return up to five chunks unless numberOfResults is set; it accepts 1 to 100.
  3. Buried chunks can go unused. Accuracy peaks with the relevant passage at the start or end of context, and is lowest in the middle.
  4. Pick k at the knee. Quality climbs, plateaus, then sags; choose the lowest k on the plateau.
  5. Retrieve wide, rerank, send few. Reranking bills per query of up to 100 chunks and the quota is 10 requests per second.
  6. Measure k on your questions. Pair real questions with their answer passages; a value copied from another system is a guess.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.