The situation
A back-office team has built a document-processing feature on Amazon Bedrock. Users upload PDFs, contracts, research reports, scanned forms, into an Amazon S3 bucket. Each one needs a generative pass: a structured summary, a set of extracted fields, and a risk classification. A single document runs from twenty seconds to four minutes through the model, depending on length, and some of the reports are two hundred pages.
The first version put the whole thing behind an API. A user uploaded through a web form, the request called Bedrock inline, and the browser waited. It worked in the demo with a two-page sample. In production it fell over immediately. API Gateway cut the integration off at its default twenty-nine seconds, so the long documents never returned, and the front-end retries re-ran the model call from scratch. A report that eventually succeeded had been summarised three or four times, and billed for every attempt. When a marketing push sent four hundred uploads in an hour, the synchronous path had no way to shed load, and half the requests errored out under Bedrock throttling.
The team now wants uploads to work at any volume, without a human watching a progress bar. The document is slow to process and expensive to process, and the user does not need the answer in the same breath as the upload.
What actually matters
The deciding property is latency tolerance. A generative pass over a long document is a batch job that happens to be triggered by a person. Nobody is staring at the screen for four minutes, and no sensible HTTP path stays open that long anyway. Once the answer is allowed to arrive minutes after the upload, the shape of the system changes. The upload becomes an event, the work becomes a queued task, and the result is written somewhere the user checks later or is notified about. Fighting to keep the call synchronous is the mistake that started this.
The second property is decoupling under bursty load. Uploads do not arrive smoothly; they arrive in clumps. Bedrock enforces per-model, per-Region quotas on your account, counted mostly in tokens per minute, and a clump goes straight through them. Something has to sit between the flood of uploads and the model and meter the work out at a rate inside those quotas. A buffer that holds pending work, and lets workers pull from it at their own pace, turns a spike into a queue that drains a little slower. Without that buffer, the spike reaches the user as errors.
The third is failure handling, and it matters more here than in a cheap CRUD system, because every retry is charged for. Model calls fail transiently, time out, or hit throttling. The naive answer, retry, is what tripled the bill in version one. Retries have to be bounded, they have to back off, and repeated failures have to land somewhere you can inspect. That means a dead-letter path for the documents that never succeed. It also means the worker has to be safe to run twice on the same document, because at-least-once delivery will hand it over twice sooner or later.
The fourth is orchestration complexity, which decides how heavy the machinery needs to be. A single summarise-and-store step is one worker. Extract text, then summarise, then classify, then write to a database, then notify, with different retry rules at each stage and a branch for documents that fail validation, is a workflow. Cramming that into one function is where worker code turns into a knot nobody can maintain. The more the stages need independent retries and visible state, the more the orchestration should be explicit rather than buried in code.
And the cross-cutting one: sometimes there is no event at all, just a pile. When the job is ten thousand documents sitting in a bucket with no deadline, a queue and a worker fleet are more machinery than the problem needs. Bedrock batch inference takes one asynchronous job, reads every record from S3, runs them, and writes the results back to S3. Select models run at 50% below on-demand inference pricing in batch. Reaching for the event-driven plumbing when a batch job would do is its own kind of over-engineering.
What we’ll filter on
- Latency tolerance, does anyone wait on the answer, or can it arrive minutes later?
- Volume and burstiness, steady trickle, spiky bursts, or a one-off pile of thousands?
- Orchestration complexity, one step, or a multi-stage pipeline with branches and per-stage retries?
- Failure handling, are retries bounded, backed off, and are dead letters captured?
- Idempotency, is a worker safe to run twice on the same document without double-charging?
- Cost shape, does the batch discount outweigh the loss of per-document immediacy?
The landscape
Synchronous request to Bedrock. The caller waits for the model inline, through API Gateway and Lambda or straight from a server. Fine for short, interactive prompts that return in a second or two. For long documents it is the anti-pattern that started this: integration timeouts, client retries that re-run expensive calls, and no way to absorb a burst. Twenty-nine seconds is the default ceiling rather than a hard wall, since Regional and private REST APIs can have it raised by quota request, and AWS may lower your Region-level throttle quota when it grants one. A raised ceiling still leaves a person watching a spinner for minutes. Rule the shape out once the work outlasts a comfortable request.
S3 event notifications as the trigger. Configure the bucket to emit an event when an object lands under a prefix, and route it onward. The upload itself becomes the signal, so there is no polling and no separate submit call. S3 sends notifications to Lambda, SQS, SNS, or Amazon EventBridge, and EventBridge is the route when you want richer routing and filtering. Delivery is designed to be at least once, usually within seconds, though it can take a minute or longer. The event carries the bucket and key, not the document, so the worker fetches the object when it runs.
Amazon SQS as the buffer. A standard queue sits between the trigger and the workers, holding pending documents so producers and consumers run at their own pace. This is what tames bursty load and what meters work against Bedrock quotas. Workers receive a message, process it, and delete it, and while they are saturated the queue simply grows. The visibility timeout hides a message while a worker holds it, and it defaults to thirty seconds. For slow model calls, set it above the worst-case processing time, or the message becomes visible again and a second worker starts the same document. Twelve hours from first receipt is the ceiling. A dead-letter queue, which must sit in the same account and Region, catches messages that fail past a set number of receives, so a poison document lands somewhere inspectable instead of cycling forever.
AWS Lambda as the worker. A function consumes from the queue, fetches the object from S3, calls Bedrock, and writes the result. Lambda scales the event source mapping with the backlog, starting at five concurrent invocations and adding up to 300 more a minute, to a ceiling of 1,250 for one SQS mapping. The constraint to respect is the fifteen-minute function timeout: comfortable for a single-document call, tight for a long multi-model chain, and a signal to split the work across steps. Functions on Lambda Managed Instances can run to ninety minutes through an event source mapping, but splitting is usually still the better answer. Capping the fan-out takes reserved concurrency on the function, or the maximum concurrency setting on the SQS event source mapping, which accepts 2 to 1,000.
AWS Step Functions for orchestration. A state machine coordinates a multi-step pipeline: extract, summarise, classify, persist, notify. Each state carries its own retry policy, catch rules, and back-off, and can branch for documents that fail a validation gate. The state of every in-flight document is visible and durable rather than implicit in a tangle of function code, and the built-in retry handling replaces plumbing you would otherwise write by hand. Standard workflows run for up to a year and are billed per state transition; Express workflows stop at five minutes, which rules them out for four-minute documents. This is the right weight when the pipeline has several stages that each need independent failure handling, and overkill for a single summarise-and-store step.
Amazon Bedrock batch inference. A single asynchronous job reads many records from an S3 input location, runs them through a model, and writes the outputs back to S3. There is no queue to run and no worker fleet to manage. A submitted job is validated, then waits in a queue until its turn comes, and it expires if it has not started before its timeout, which is set between 24 and 168 hours. Batch does not support tool calling or structured output, so field extraction has to be prompted rather than schema-enforced. This is the fit for high-volume, latency-tolerant work: a nightly enrichment of a whole table, a one-off pass over an archive. It is the wrong tool when documents arrive one at a time and each needs a timely answer.
Evaluation
Side by side
| Building block | Latency fit | Volume fit | Orchestration | Failure handling | Cost shape |
|---|---|---|---|---|---|
| Synchronous to Bedrock | Interactive only | Low, no burst absorption | ✗ | Client retries re-run work | Pay per attempt |
| S3 event notification | Fires on upload | Any | ✗ (just the trigger) | Hands off to target | Negligible |
| SQS buffer + DLQ | Seconds to minutes | ✓ Absorbs bursts | ✗ | ✓ Bounded retries, DLQ | Cheap per message |
| Lambda worker | Minutes (15-min timeout) | ✓ Scales to 1,250 | Single step | Redelivery after timeout | Pay per run, idle-free |
| Step Functions | Minutes to hours | ✓ | ✓ Multi-stage, branches | ✓ Per-state retry and catch | Pay per transition |
| Bedrock batch inference | Not per-upload | ✓ High volume | Single bulk job | Per-record error counts | 50% below on-demand |
Read the table against the team’s feature. The upload fires an S3 event, SQS buffers the burst with a DLQ behind it and a visibility timeout tuned to the four-minute worst case. The choice between a lone Lambda and a Step Functions pipeline comes down to whether the extract-summarise-classify-persist-notify chain needs independent per-stage retries, which it does. The nightly enrichment of the back catalogue is the one piece that suits batch inference instead, because it is a pile with no deadline.
The solution
For the live upload path, the spine is S3 notification into SQS into Lambda, and the queue settings make or break it. The visibility timeout has to exceed the worst-case processing time plus a margin. With documents that can take four minutes, a timeout of six keeps the message hidden while a worker grinds through a two-hundred-page report. Set it too short and a second worker starts the same document, which on Bedrock means paying twice. The maxReceiveCount in the redrive policy bounds how many times a failing message is retried before it moves to the dead-letter queue. A corrupt PDF or an unsupported format then lands in the DLQ after a few attempts instead of blocking the queue or looping forever. Give the DLQ a longer retention period than the source queue, because a standard-queue message keeps its original enqueue timestamp when it moves.
Idempotency is the piece people skip and regret. S3 notifications, SQS and Lambda event source mappings all deliver at least once, so the same document will occasionally arrive twice. Every retry, whether from a short visibility timeout, a redrive, or a transient error, is a chance to run the expensive model call again. The defence is a deterministic result key, typically derived from the object key and version or a content hash, and a check before work. If the result for this document already exists in the output store, the worker returns without calling the model. A duplicate delivery then ends in an S3 lookup rather than a second Bedrock charge, which is what lets you set generous retry policies.
Concurrency against Bedrock quotas is the other tuning knob. One SQS event source mapping scales to 1,250 concurrent invocations, and Bedrock throttles calls past your per-model quota, so an uncapped worker fleet converts a backlog into a wall of throttling errors. The maximum concurrency setting on the event source mapping caps the fan-out at a number the quota sustains, anywhere from 2 to 1,000. Reserved concurrency on the function does a similar job at function level, and AWS advises keeping it at or above the mapping’s maximum. The queue absorbs the rest and drains at that steady rate. Pair the cap with retry-on-throttle and a short back-off, so an occasional throttled call recovers rather than falling through to the DLQ. Requesting a quota increase is the move when the sustained rate genuinely needs to be higher.
Step Functions fits once the work is a pipeline rather than a step. Extract text, summarise, classify, write to the database, send the notification: each of those can fail independently, and each needs its own retry and catch behaviour. A state machine gives durable, inspectable state for every document, instead of a mega-function that logs its own errors and carries on. A Catch on a state routes a failed document to a handling branch, a Retry block applies bounded exponential back-off per state, and the execution history shows where a given document is or why it stopped. Standard workflows are billed per state transition and the definition is yours to maintain, so a genuine single-step job does not need one. Start with the Lambda-off-a-queue shape, and graduate when the stages and their independent failure handling actually appear.
Batch inference removes the plumbing entirely, for the workloads that suit it. When the job is enriching a whole table overnight, or summarising an archive of ten thousand filings with no per-item deadline, one asynchronous S3-to-S3 job at half the on-demand token price beats building and running a queue and a worker fleet. What you give up is immediacy: the job validates, waits its turn, and completes on its own timeline. Track it with the record counters on GetModelInvocationJob, or take an EventBridge notification on the job state change instead of polling. Many real systems run both, the event-driven path for live uploads and a nightly batch job for the backlog.
Worked example
A user drops contracts/2026/acme-msa.pdf into the ingest bucket. The flow that follows:
S3 (ObjectCreated on contracts/*)
-> event notification
SQS ingest-queue (visibility timeout 360s; DLQ after 3 receives)
-> Lambda event source mapping (maximum concurrency 20)
Lambda worker
1. derive result key = sha256(bucket, key, versionId)
2. if result exists in results bucket -> return, no model call
3. fetch object from S3
4. call Bedrock (retry on throttling, short back-off)
5. write summary + fields + classification to results bucket
6. return cleanly; Lambda deletes the batch from the queue
Failure past 3 receives -> ingest-dlq (inspect corrupt / unsupported docs)
Lambda deletes a batch from the queue only when the invocation returns successfully. A worker that crashes at step 4 leaves the message to reappear after the visibility timeout, and three failed receives send it to the DLQ. With more than one message in a batch, report partial batch failures in the response, so the messages that succeeded are not redelivered with the one that failed. The idempotency check at step 2 turns a duplicate delivery into an S3 lookup rather than a second Bedrock charge. A maximum concurrency of twenty holds the worker fan-out under the model’s quota, so a four-hundred-upload burst becomes a queue that drains twenty-wide.
When the same team later needs extract, then summarise, then classify, then persist, then notify, each with its own retry rules and a branch for documents that fail validation, the single worker becomes a Standard state machine triggered off the same queue. Each stage gets its own Retry and Catch. The untouched back catalogue of fifty thousand old contracts, with no deadline on it, goes through one Bedrock batch inference job reading from S3 and writing back to S3 at half the price, rather than through the live queue.
What’s worth remembering
- A long generative pass over a document is a batch job triggered by a person; once the answer may arrive minutes later, the synchronous constraint disappears and the design gets simpler.
- Put SQS between the trigger and the workers to absorb bursty uploads and meter work against Bedrock’s per-model, per-Region quotas, and set the visibility timeout above the worst-case processing time.
- Make the worker idempotent with a deterministic result key and a check before work, because S3 notifications, SQS and Lambda event source mappings each deliver at least once.
- Cap worker fan-out with the maximum concurrency setting on the SQS event source mapping, which takes 2 to 1,000, and pair the cap with retry-on-throttle and back-off.
- A high-volume pile with no per-item deadline suits Bedrock batch inference: one asynchronous S3-to-S3 job at 50% below on-demand, waiting its turn in a queue rather than answering per upload.