The situation
A support-automation team runs a customer-managed Amazon Bedrock knowledge base over roughly forty thousand documents in S3: product manuals, pricing sheets, policy pages, and a large archive of resolved support tickets. An agent retrieves the top passages for each customer question and grounds its answer in them. When the knowledge base was first built, every document was embedded once and written to the vector store, and retrieval has worked well since.
The problem is keeping it current. Pricing sheets change weekly, policy pages change a few times a month, and the manual archive barely moves. To stay fresh, the team wired a nightly job that re-ingests the entire data source. It works, but the embedding cost of pushing forty thousand documents through the model every night now dwarfs the cost of the queries the knowledge base actually answers, and the sync is still running well into the morning. Worse, when a document is deleted at source, nobody is sure the old passages ever leave the index, so retrieval sometimes surfaces a policy that was retired months ago.
The team wants fresh answers without paying to re-embed a corpus that mostly did not change, and without stale passages lingering to outrank the current ones. Underneath the nightly-cost complaint are three separate questions: what has to be re-embedded, when the work should run, and which facts do not belong in a knowledge base at all.
What actually matters
Embedding is the expensive, slow part of ingestion, and it scales with how much text you push through the model, not with how much of it changed. Every document that gets re-processed is parsed, chunked, embedded, and written to the vector store, and the embedding-model invocation is where both the money and the minutes go. Re-embedding forty thousand documents to reflect a change in forty of them means paying for the other thirty-nine thousand nine hundred and sixty for nothing. Full re-ingestion of a large corpus is a cost you almost never need to pay.
Bedrock knowledge bases already work this way. After the first sync of a data source, every later sync is incremental: the connector compares the current state of the source against what it has already indexed, and re-processes only the documents added, modified, or deleted since. An unchanged document is skipped rather than re-parsed, re-chunked and re-embedded. A sync after forty documents moved therefore does forty documents of embedding work, not forty thousand. The StartIngestionJob response reports the split directly, in numberOfDocumentsScanned against numberOfDocumentsSkipped, numberOfNewDocumentsIndexed, numberOfModifiedDocumentsIndexed and numberOfDocumentsDeleted. The nightly full re-ingest overrides that behaviour rather than using it, most likely by rebuilding the data source instead of letting the connector diff.
Given incremental sync exists, the lever is when it runs. A customer-managed knowledge base has no built-in schedule, so something outside it has to call StartIngestionJob. A schedule is the simple shape: an EventBridge Scheduler universal target invokes that API every few hours, and each run’s embedding cost is bounded by how much changed since the last one. An event-driven trigger, where an S3 event notification starts a sync as an object lands, makes answers current within minutes of a change, and it needs more plumbing to keep the job count down. One constraint shapes both. Bedrock allows one concurrent ingestion job per data source and one per knowledge base, so a StartIngestionJob sent while a job is running returns a ConflictException. Sync frequency is capped by how long a sync takes, whatever starts it.
Deletes and updates are where staleness does real damage, because a stale passage that retrieval ranks above the current one produces a wrong answer rather than a missing one. When a document changes, its old ChunkingSplitting documents into retrievable pieces before embedding them – small enough to match precisely, big enough to still make sense. have to leave the index. Incremental sync does that when the S3 objects are what the connector reads: modify an object and its old vectors are replaced, delete an object and its vectors are removed.
The failures are the changes a sync cannot observe. Narrowing the data source’s inclusion prefix, or moving an object to a key outside it, takes the document out of what the connector crawls without recording a deletion, so its chunks stay in the vector store and stay retrievable. Lifecycle expiration is the same shape from S3’s side: s3:ObjectRemoved:* notifications do not fire for objects that a lifecycle configuration deleted, which emits s3:LifecycleExpiration:* instead, so a pipeline wired only to ObjectRemoved never starts a sync for those. Direct ingestion cuts the other way. IngestKnowledgeBaseDocuments writes into the vector store without touching the bucket, so DeleteKnowledgeBaseDocuments leaves the S3 object in place and the next sync re-indexes it.
Metadata is the query-time lever for the versions that legitimately coexist. Bedrock reads a filename.extension.metadata.json file stored beside each S3 object, under 10 KB, with attributes typed STRING, NUMBER, BOOLEAN or STRING_LIST. A filter on a Retrieve or RetrieveAndGenerate call tests them. That filter includes and excludes; it does not reorder, so it cannot make a recent chunk outrank an older one. What it can do is drop the older one from the candidate set before ranking, with greaterThan on an effective date. Those numeric comparisons accept numbers only, so store the date as 20260731 or as epoch seconds rather than a date string. Reranking is the separate feature that reorders results already returned.
Facts that are live lookups of a value do not belong in a knowledge base at any refresh rate. A current account balance, today’s inventory count, a live order status, a price that changes intraday: these are not documents to embed, they are values to read. Retrieval answers as of the last sync, so for a fact that must be correct to the second, no sync cadence is fast enough and every sync is wasted embedding. Those belong in a tool call, where the agent queries the system of record at question time.
What we’ll filter on
- Change rate, how often the source documents actually change, and what fraction of the corpus moves per period.
- Freshness requirement, how stale an answer is allowed to be before it is wrong, per document type.
- Re-embedding budget, how much embedding cost and sync latency the corpus size implies per full pass.
- Delete and update fidelity, whether retired content reliably leaves the index rather than lingering.
- Trigger fit, whether a schedule or an event-driven sync matches the freshness requirement without over-syncing.
- Fact volatility, whether the value is a document to retrieve or a live reading to look up at query time.
The landscape
Full re-ingestion. Clear the vectors and re-embed the whole data source. The only case this is worth doing is a genuine reset: a new embedding model, a chunking-strategy change that invalidates every existing vector, or a first build. As a routine freshness mechanism on a large, slow-changing corpus it is the most expensive option by a wide margin, because you pay to re-embed everything to reflect a change in almost nothing.
Incremental sync (the default behaviour). After the first sync, the Bedrock data-source connector diffs the source and re-processes only added, modified, and deleted documents. This is the baseline to be on: the cost of a sync tracks the volume of change, not the size of the corpus. Deletes and updates are part of the diff, so retired content leaves the index when the connector sees the deletion. The work is to trigger it well, not to replace it.
Scheduled incremental sync. Start an ingestion job on a fixed cadence. On a customer-managed knowledge base there is no built-in schedule, so an EventBridge Scheduler universal target calls StartIngestionJob at an interval you set. Freshness is capped at the interval length, so a weekly pricing change on a daily schedule can be up to a day stale. A Bedrock Managed knowledge base does have a built-in sync schedule, but its options are daily, weekly and monthly, which is too coarse for pages that move hourly.
Event-driven incremental sync. An S3 event notification fires a Lambda that starts a sync when an object is created or removed. Answers go current within minutes of a change, which suits documents that must not lag. The costs are operational: S3 notifications arrive at least once and in no guaranteed order, so a burst produces duplicate and out-of-order triggers, and only one ingestion job can run per data source at a time. Best reserved for the subset of the corpus that needs minute-level freshness.
Metadata-filtered retrieval. Attach a numeric effective date to each document and filter queries on it, so retrieval excludes anything superseded or outside the window you care about. This is a query-time lever, not an ingestion one. It removes older chunks from the candidate set rather than deleting them, and it ranks nothing, so it pairs with any of the sync options above rather than replacing them.
Live data or tool call. For volatile facts, skip retrieval and have the agent call the system of record at question time, reading the current balance, price, or status directly. Correct to the moment, no embedding, no sync to keep current. It fits facts that are values rather than passages of prose, and it needs the tool and permissions wired up, but for the real-time slice it is the only mechanism that is ever fresh.
Evaluation
Side by side
| Mechanism | Cost shape | Freshness | Handles deletes | Best for |
|---|---|---|---|---|
| Full re-ingestion | Whole corpus every run | As of last full pass | ✓ | Model or chunking change, first build |
| Incremental sync | Only changed documents | As of last sync | ✓ | The default for any changing corpus |
| Scheduled incremental | Change-per-interval | Capped at interval | ✓ | Steady, predictable change rates |
| Event-driven incremental | Change-per-event | Minutes | ✓ | Documents that must not lag |
| Metadata-filtered retrieval | Query-time only | Excludes superseded | ✗ (query-time, not ingestion) | Old and new legitimately coexist |
| Live data / tool call | Per query, no embedding | Real time | Not applicable | Volatile values, not documents |
Reading the table against the corpus: the manual archive needs only a slow scheduled incremental sync; the pricing and policy pages call for event-driven incremental so a change is reflected in minutes; every document type benefits from a numeric effective date so a query can exclude superseded versions; and any genuinely live fact (a customer’s current plan status, today’s price) belongs in a tool call, not the index at all. The nightly full re-ingest serves none of these well.
The solution
Getting off full re-ingestion is the first and largest change, and it is mostly a matter of stopping the wrong thing. The nightly job is almost certainly rebuilding rather than diffing, either by recreating the data source or clearing vectors before ingesting. Run StartIngestionJob against the existing data source instead, and the connector re-processes only what changed since the last successful sync. The same forty-document change that took a full corpus of embedding overnight becomes forty documents of work, and the sync finishes in a fraction of the time.
Choosing the trigger is where the freshness-versus-cost trade gets made, and it is worth making per document type rather than once for the whole knowledge base. The slow-moving manual archive does not justify event plumbing; a scheduled incremental sync every few hours, or even daily, keeps it current enough and keeps the job count low. The pricing and policy pages are the opposite: a stale price is a wrong answer with real consequences, so an S3 event notification firing a Lambda that starts a sync is worth the extra machinery. The thing to get right in the event-driven path is coalescing. Only one ingestion job runs per data source at a time, so a bulk update that rewrites two hundred objects would otherwise produce one job and a hundred and ninety-nine ConflictExceptions. Buffer the notifications in a short SQS-backed window and start a single job for the batch. Passing a clientToken gives a second line of defence, since Bedrock ignores a repeat of a token it has already seen.
Deletes deserve explicit attention because they fail without an error. As long as the S3 objects are what the connector reads, removing one and running a sync removes its vectors, and a modified object replaces its old chunks. The retired policy the team keeps seeing points at deletions that never reached a sync: an object aged out by a lifecycle rule, a prefix narrowed, or a document removed through DeleteKnowledgeBaseDocuments while the S3 object stayed. Make the bucket authoritative and change content only by changing objects in it, so every add, update, and delete flows through the same diff. Set the data source’s data deletion policy to DELETE as well, so tearing the data source down clears its vectors instead of retaining them.
Metadata is the guard for the versions that stay. Write a filename.extension.metadata.json beside each document carrying a NUMBER effective date, and pass a greaterThanOrEquals filter on every retrieval call so superseded versions never reach the candidate set. This does not substitute for deleting stale chunks, and it does not reorder anything; it removes documents a ranking pass would otherwise have to choose between. Editing one of those files is also unusually cheap: when only .metadata.json files changed, the content is not a CSV, and the data source has no custom transformation Lambda, Bedrock merges the new metadata into the stored vectors without calling the embedding model at all.
The last pick is a boundary, not a mechanism. For facts that change faster than any reasonable sync (a customer’s live plan status, an intraday price, a current stock level) retrieval is the wrong tool, because it answers as of the last sync and every sync spends embedding on a value that will be wrong again by lunchtime. Wire those as a tool the agent calls against the system of record at question time. The knowledge base then holds the durable prose (how cancellation works, what the tiers include) while the volatile numbers come from a live call. This mirrors the reasoning behind reaching for a tool call rather than the model’s own weights when the answer depends on current data.
Worked example
A pricing sheet is updated in S3 at 09:00 on a Tuesday. Under the nightly full re-ingest, the change is invisible until the next run at 02:00 Wednesday, so for seventeen hours the assistant quotes the old price, and the run that finally picks it up re-embeds all forty thousand documents to reflect one changed file.
Reworked, the same change flows differently. The pricing prefix in S3 has event notifications enabled; the 09:00 write lands on an SQS queue, and a Lambda drains it after a short window in case more sheets follow, then calls StartIngestionJob against the existing data source. The connector diffs the source, finds one modified object, re-embeds that sheet’s chunks, replaces the old vectors, and finishes in seconds. By 09:03 retrieval returns the new price. The sheet also carries an effective date, so a retrieval filter excludes the superseded chunk even in the window before its vectors are replaced.
# S3-notification-driven sync (conceptual)
on s3:ObjectCreated:* | s3:ObjectRemoved:* for prefix pricing/:
buffer notifications on SQS for 60s # coalesce a burst into one job
aws bedrock-agent start-ingestion-job \
--knowledge-base-id ${KB_ID} \
--data-source-id ${DS_ID} \
--client-token ${BATCH_TOKEN}
# connector re-processes only changed/added/deleted objects
Note the event filter. Lifecycle expiry would not appear here, so any pricing sheet aged out by a lifecycle rule needs s3:LifecycleExpiration:* wired up too, or its chunks stay. And the fact that should never have been a document, the customer’s current plan and next billing date, is not retrieved at all: the agent calls a billing tool at question time and reads it live. The pricing prose is fresh within minutes for one document’s worth of embedding, the account fact is correct to the second with none. The seventeen-hour lag and the nightly forty-thousand-document bill are both gone.
What’s worth remembering
- Embedding cost scales with the volume of text re-processed, not with how much changed, so full re-ingestion of a large corpus is a bill you almost never need to pay.
- After the first sync, Bedrock knowledge base syncs are incremental: only added, modified, and deleted documents are re-parsed, re-chunked and re-embedded, and unchanged ones are skipped.
- A customer-managed knowledge base has no built-in schedule, and only one ingestion job runs per data source at a time, so coalesce bursts of S3 notifications into one job rather than firing one per object.
- Metadata filters include and exclude, they never reorder, and
greaterThantakes numbers only, so store an effective date as a number and use it to drop superseded chunks before ranking. - Deletes fail without an error when the sync cannot see them: a lifecycle expiry, a narrowed inclusion prefix, or a direct-ingestion delete that left the S3 object in place all leave chunks behind.
- For facts that change faster than any sync (live status, intraday price, current balance) retrieval is the wrong tool; call the system of record at question time and keep only the durable prose in the index.