Exam Room · Advanced Generative AI Developer

How to Build a Citations-Required RAG Over 50K Internal Documents

· 29 min read

Generative AI Development · part of The Exam Room

The situation

A 6,000-person enterprise is standing up an internal assistant. The corpus is ~50,000 documents across four domains: HR policies, engineering runbooks, security guidelines, and product specs, totalling ~5 GB of mostly text-dense PDFs, Markdown, Word, and Confluence exports. New documents land weekly, old ones get superseded, a handful are retracted. The assistant has to reflect the current state within a day of a change.

On the answer path:

  • P95 end-to-end latency < 3 s from question to last TokenThe unit of text an LLM actually sees – usually a short character sequence, not a whole word., across retrieval, generation, and network.
  • Document-level access control. An engineer asking “what are the band-5 engineering salaries?” must get a decline, not an HR document. A security auditor asking about an incident-response runbook gets the runbook. Identity drives what the retriever can see.
  • Citations on every answer. Every factual claim points back to a source chunk. No citation, no answer.

What actually matters

Identity meeting retrieval is where a system like this succeeds or fails, so the first question is who owns that boundary? A product team that ships “the assistant” without owning the access-control fabric under it is building a compliance incident with a generative front-end. The design has to make the seam explicit: identity in, filter out, retriever sees only what the caller is allowed to see. Filtering results after retrieval leaves the top-K polluted with chunks the user can’t read, and filtering at generation leaves the citation hanging off something the user shouldn’t have seen in the first place.

The second is what’s the blast radius of a bad answer? An engineer who asks about someone else’s salary and gets a decline is fine. An engineer who asks about someone else’s salary and gets the answer is a wrongful-disclosure incident, and the remediation runs to legal notice, HR escalation, and a six-month trust deficit with the workforce that was just asked to share more data with the tool. A single leakage outweighs every other risk on the project. That shape pushes the design toward managed components where the access-control path is a first-class API rather than glue the team maintains.

The third is what happens as the corpus grows? Five gigabytes today, seven next year, thirty when the internal wiki finally gets ingested. The ingestion story has to be incremental by default, since a full weekly reprocess of 5 GB is doable and a full weekly reprocess of 30 GB runs for hours. The VectorAn ordered list of numbers – in AI usage, almost always an embedding – and by extension the databases that index them for nearest-neighbour search. bill scales with vector dimensions × chunks × replicas, so the EmbeddingA fixed-length vector of floats that represents a piece of text (or image, or other thing) in a space where similar meanings sit close together. model choice sets a multi-year storage footprint. Changing embedding models means reindexing everything, so the dimension trade-off is quick to set at install time and slow to undo.

The fourth is what are the failure modes we have to design against? A citation the user can’t load because the S3 object is gated by a different policy. A retrieval that returns zero chunks for a legitimate question because the filter is too tight. A chunking strategy that slices a procedure in half, so the generated answer joins two halves of two runbooks. A metadata-sidecar path where a file was added without its .metadata.json and therefore has no allowed_groups, defaulting to nobody or everybody depending on how the filter is composed. Each of those needs a test, a runbook, and a monitoring line. The managed service handles about half; the application team owns the rest.

The fifth is where does a small platform team want to spend its operational attention? Not on owning a vector-store operator, not on writing chunking pipelines, not on re-implementing citation extraction for the fourth time. Managed services remove that work and some flexibility with it; the trade works when the workload is standard and fails when it has an unusual shape. A 50K-document corpus with vanilla group-based access control is standard. A SOX-grade audit requirement with multi-hop ACL joins is not, and calls for SQL.

Finally: what does “current state” mean in practice? The brief says “within a day” but the business will discover it means “within an hour” the first time a retracted policy still turns up in an answer. The ingestion cadence has to scale from weekly-cron down to per-object event without re-architecting, because the product requirement will tighten under production pressure.

What we’ll filter on

Five filters, and the landscape either clears them or doesn’t.

  1. Document-level access control enforced during retrieval. Not a post-hoc scrub of results, otherwise the top-K is polluted with chunks the user can’t see and quality collapses.
  2. Sub-3-second end-to-end latency at P95. Retrieval under a second, generation streamed, first tokens visible to the user inside one.
  3. Citations that survive the model summarising or paraphrasing. The generation path has to propagate “which chunk came from which document” all the way to the response.
  4. Incremental weekly ingestion. New files picked up, changed files re-embedded, deleted files removed. Not a full weekly reprocess of 5 GB.
  5. Reasonable operational overhead. A small platform team. Managed components where the differentiation isn’t worth hand-rolling.

The landscape

Five plausible shapes on AWS.

Fine-tune a foundation model on the corpus. No retrieval at all, the knowledge goes into the weights. Weekly refresh means weekly fine-tune cycles at 5-GB scale. Citations are impossible because fine-tuning merges sources into weights with no pointer back. Per-user access control is impossible because once a chunk is in the weights, every user sees it.

Bedrock Knowledge Bases, in its customer-managed form. A managed RAG pipeline that ingests documents from a data source (S3, SharePoint, Confluence, Salesforce, web crawler, custom), chunks them, embeds them through a chosen model, stores the vectors in a store you provision, and exposes two runtime APIs: Retrieve for raw chunks and RetrieveAndGenerate for the full round-trip with citations. Eight supported vector stores: OpenSearch Serverless, OpenSearch managed clusters, S3 Vectors, Aurora pgvector, Neptune Analytics (GraphRAG), Pinecone, Redis Enterprise Cloud, MongoDB Atlas. Text embedding models are Titan Embeddings G1 (1,536 dim), Titan Text Embeddings V2 (256 / 512 / 1,024), and Cohere Embed English and Multilingual v3 (1,024 each), with multimodal models alongside them at 1,024. Metadata filtering during retrieval and citations in generation are first-class.

Custom RAG with Bedrock + OpenSearch Serverless vector engine. Same substrate as the most common Knowledge Bases configuration, but you write the pipeline: ingestion Lambdas, embedding invocations, k-NNThe retrieval question itself: given a query vector, return the k closest vectors under the index’s distance metric – answered exactly by comparing against everything, or quickly by an ANN index. mappings, prompt assembly, citation extraction. Every component is under your control and yours to operate. OpenSearch Serverless supports HNSW with Faiss, Euclidean / cosine / dot-product metrics, and up to 16,000 dimensions. You set a minimum and maximum OCU count for indexing and search separately, from zero upwards, at USD$0.24 per OCU-hour.

Custom RAG with Bedrock + Aurora PostgreSQL pgvector. Same DIY pipeline, but the vector store is Aurora with pgvector 0.5.0+ and HNSW indexes on a vector(n) column. Knowledge Bases can also consume Aurora as a vector store via the RDS Data API plus Secrets Manager. The selling point is SQL: embeddings sit next to the metadata you already keep relationally, and filters become ordinary WHERE clauses.

Custom RAG with Bedrock + Amazon Kendra. Kendra is an intelligent search service rather than a vector database, with its own ranking models and built-in document-level security, and it was a credible retrieval layer for exactly this shape of problem. AWS put it into maintenance mode on 30 June 2026 and closed it to new customers on 30 July 2026, so a new build cannot pick it, and it is left out of the comparison below.

Evaluation

Side by side

Option Access control in retrieval <3 s P95 Citations Incremental sync Low ops overhead
Fine-tune foundation model ✗ ✓ ✗ ✗ ✗
Bedrock Knowledge Bases ✓ ✓ ✓ ✓ ✓
Custom RAG on OpenSearch Serverless ✓ ✓ ✓ ✓ ✗
Custom RAG on Aurora pgvector ✓ ✓ ✓ ✓ ✗

Matching the shape to the managed service

Authenticated user engineer, on-call session from IdP Identity translation server-side only groups = [engineering, on-call] Filter composition orAll listContains allowed_groups ∈ user groups Bedrock Knowledge Bases Retrieve + metadata filter HNSW cosine, numberOfResults 10 hierarchical: child 300 tok, parent 1,500 tok returned filter applied *during* k-NN, not after HR chunks never enter top-K OpenSearch Serverless Titan V2 1,024-dim metadata sidecars autoscaling OCUs, HNSW + Faiss Weekly ingestion StartIngestionJob on S3 deltas only, per-object triggers ready Claude Sonnet RetrieveAndGenerate $output_format_instructions$ preserved Answer with inline citations each span linked to retrievedReferences[*].location.s3Location
Identity in, filter composed server-side, metadata filter applied during retrieval (green dashed), citations emitted by preserving the default prompt template's `$output_format_instructions$` placeholder.

The solution

Bedrock Knowledge Bases on OpenSearch Serverless, with Titan Text Embeddings V2 at 1,024 dimensions and every retrieval filtered by the caller’s group membership.

Chunking. Five strategies: default (~300 tokens, sentence-aware), fixed-size (tunable), hierarchical (child for precision, parent for context), semantic (LLM-driven boundaries with buffer and percentile threshold), no-chunking (one chunk per document, which gives up page numbers in citations). For runbooks and policies, structured documents where the answer is a two-sentence span but the generator needs surrounding subsection context, hierarchical is the better fit. Child 300 tokens, parent 1,500. Above 8,000 combined tokens you can exceed metadata-size limits, and AWS does not recommend hierarchical chunking on an S3 vector bucket at all.

Embedding model. Titan Text Embeddings V2 at 1,024 dimensions suits an English corpus with a moderate per-vector footprint. Dropping to 512 halves vector storage at some retrieval-quality cost. Cohere Embed English v3 is the alternative at the same 1,024 dimensions, and its multilingual sibling is the choice once the corpus stops being all English. Dimensions are locked to the embedding model, so switching models means reindexing the whole corpus.

Access control through metadata filtering. Every document has a companion <filename>.metadata.json declaring allowed_groups, domain, classification, effective_date. Every retrieval call passes a filter composed server-side from the authenticated caller’s group membership:

{
  "vectorSearchConfiguration": {
    "numberOfResults": 10,
    "filter": {
      "orAll": [
        {"listContains": {"key": "allowed_groups", "value": "engineering"}},
        {"listContains": {"key": "allowed_groups", "value": "on-call"}}
      ]
    }
  }
}

The filter is applied during vector search, not after. Chunks whose metadata doesn’t satisfy it never enter the top-K. Available operators: equals, notEquals, greaterThan(OrEquals), lessThan(OrEquals), in, notIn, startsWith (OpenSearch Serverless only), stringContains, listContains, and andAll / orAll to combine 2 to 5 of them. numberOfResults runs from 1 to 100 and defaults to 5. Enough for group-based rules; not enough for full ABAC with clearance-level comparisons.

The filter is composed by a trusted backend on every call. If the browser gets to construct it, there’s no access control at all.

Incremental ingestion. StartIngestionJob syncs the data source incrementally: an unchanged document is skipped, one whose content or metadata changed is re-parsed, re-chunked, re-embedded and re-indexed, a new one is ingested, a deleted one is removed from the vector store. Weekly cron via EventBridge; per-object triggers from S3 event notifications when the product tightens to near-real-time.

Citations. RetrieveAndGenerate returns a citations array linking spans of the generated text to retrieved chunks plus their S3 URIs and metadata. Citations require the $output_format_instructions$ placeholder in the prompt template; remove it to hand-tune instructions and the citations stop appearing, with no error to say so.

When Aurora pgvector is the better pick instead

Reach for Aurora pgvector directly when the access-control logic exceeds what metadata-filter operators express: multi-hop joins across user / group / ACL / classification tables, clearance-level ≤ user-clearance via a lookup table, time-windowed validity (effective_date <= now() AND (expiry_date IS NULL OR expiry_date > now())). SQL handles all of that; metadata attributes can’t. It is also the right call when the team already runs Postgres, so adding pgvector 0.5.0 or later plus an HNSW index is a smaller jump than owning an OpenSearch Serverless collection, or when transactional consistency between documents and metadata matters (an ACL change and its embedding update atomically, no stale-filter window).

For 50,000 documents with a vanilla group-membership filter, Aurora is overkill. For 5 million documents with SOX-grade audit against a mature Postgres estate, it is the shape that fits.

When a managed knowledge base removes the ACL subsystem

This design builds entitlement out of metadata filters, and one shape provides it instead. A Bedrock Managed Knowledge Base, where Bedrock runs the vector store as well as the pipeline, ingests source permissions through its connectors and filters retrieval on a userContext you pass, so document-level access control comes with the service rather than out of sidecars you maintain. The constraints are real: a custom embedding model has to be float32 at 1,024 dimensions, chunking is built-in or fixed-size, search is always hybrid, and there are seven connectors (S3, SharePoint, Confluence, Google Drive, OneDrive, web crawler, custom) rather than any source you can write an ingestion Lambda for.

Where the entitlement rules are ordinary group membership and the content sits in supported sources, the managed shape removes a whole subsystem and is the better answer. Where they are the multi-hop, time-windowed rules described above, the filters (or Postgres) still are.

ACL-aware retrieval fails closed. userId is the user’s email address in the underlying source, and a request that omits userContext gets zero results from every ACL-enabled data source; a document whose connector produced no ACL goes to nobody. The gap is a non-ACL data source in the same knowledge base, which returns its documents to every caller whatever user context is passed. Keep the ACL flag on every source carrying anything restricted, and hold it with a test that asserts a document the caller should not see does not come back.

Worked example

One question, end to end. An engineer asks “What’s the runbook for rotating the production database password?” Groups ["engineering", "on-call"].

  1. Identity translation. Backend looks up groups, confirms the session is live, composes the retrieval filter.
  2. Embed the query. Titan V2 returns a 1,024-dim vector in ~30-80 ms.
  3. Vector search with filter. Retrieve with numberOfResults: 10 and the orAll filter. OpenSearch Serverless runs HNSW k-NN with metadata filtering during search, returning ten chunks. HR chunks never contribute noise. ~100-250 ms.
  4. Hierarchical replacement. Child chunks sharing a parent collapse to the parent. Ten children might become six parents, each 1,500-token, each with surrounding procedural context.
  5. Prompt assembly. Knowledge Bases fills the default prompt template, $output_format_instructions$ included; drop that placeholder and citations stop.
  6. Generation. RetrieveAndGenerate calls Claude Sonnet via a cross-region inference profile. First token ~800 ms; a 300-token answer finishes in ~1.8 s.
  7. Citations. Response includes a citations array linking spans of generated text to retrieved chunks plus S3 URIs. The app renders each as a numbered inline reference.

Total end-to-end: embedding 60 ms + vector search 180 ms + orchestration 50 ms + first-token 800 ms + streaming 1,000 ms = ~2.1 s P95. Generation dominates; retrieval barely registers.

What’s worth remembering

  1. Knowledge Bases manages RAG. Retrieve returns raw chunks; RetrieveAndGenerate runs the full round-trip with citations.
  2. Chunk hierarchically for structured documents. Child chunks match precisely; their parent chunks give the generator surrounding context.
  3. Filter during search, not after. Metadata filters keep disallowed chunks out of the top-K, so they never reach the generator.
  4. Compose filters server-side. A trusted backend translates identity to groups on every call; the browser never builds the filter.
  5. Citations need $output_format_instructions$. Remove it from the prompt template and citations vanish, with no error.
  6. Managed Knowledge Bases filter on userContext. Permissions come from connectors; the limits are 1,024-dimension embeddings, hybrid-only search and seven connectors.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.