The situation
An overnight job takes a batch of a few thousand supplier documents, extracts structured fields from each with a foundation model, validates the extraction against a schema, writes the clean records to a data store, and, when a document fails validation twice, routes it to a human reviewer before carrying on. Some of those steps are a model call. Most of them are not. The whole thing has to run to completion even when a single model invocation throttles, a Lambda times out, or a reviewer takes two days to respond, and the operations team has to be able to open a run afterwards and see exactly which documents took which path.
The first instinct is to reach for a Bedrock agent, because the model is doing the interesting work. That instinct is worth questioning. The model is one participant in a workflow that is mostly known in advance, mostly deterministic, and mostly about moving data reliably between AWS services with retries and a human pause in the middle.
Three engines could carry it: a model-driven Bedrock agent, a visually-defined Bedrock Flow, and an AWS Step Functions state machine. They differ in where the decision about what happens next lives, and they differ far more in how each one behaves on a run that has to survive the night.
What actually matters
Who decides the control flow comes first, and for this job it settles itself. Extract, validate, store, escalate a second failure to a reviewer: a person can draw that in a minute, and it is the same shape for every document, so there is nothing here for a model to discover at run time. The case for drawing a known sequence rather than having the model rediscover it on every run is made at length for a pipeline whose graph a designer draws, and none of it changes for a batch job. What is left is harder: which runtime should own the sequence once it is drawn.
Durability settles most of it. A batch running for hours across thousands of documents will meet throttling, transient errors, and slow dependencies, and it has to survive all of them without a process of yours held open for the duration. Step Functions keeps execution state on the service side. A Standard execution outlives whatever started it, runs for up to a year, carries its own error handling per state, and can sit idle for two days waiting on a reviewer. Neither an agent invocation nor a Flow invocation runs on that clock, and adding it around either one means writing and operating a workflow engine of your own.
Then the shape of the work. Most of this job is not model calls. It is reads and writes to a data store, schema validation, fan-out across many records, and a human approval pause, with a Bedrock call inside one step of each item. Step Functions reaches over two hundred AWS services through its SDK integrations, fans work out with its Map state, runs branches in parallel, waits on a callback token for a human decision, and calls Bedrock through an optimised integration where the model is genuinely needed. When the generative work is one participant in a longer business process, the orchestrator that reaches furthest across the rest of AWS should own the sequence.
Last, what the run has to prove afterwards. The operations team needs to open last night’s run and see which document took which path, which was escalated, which field was overwritten. A Standard execution keeps a history you can inspect state by state, and Step Functions retains it for 90 days after the run closes, so the record is a property of the service rather than something you assemble from logs. Model-chosen control flow leaves a reasoning trace instead, which is thin evidence when what you have to show is that a required step ran on every document.
The deciding line, then, is whether this is a generative pipeline that happens to branch or a durable business process that happens to call a model.
What we’ll filter on
- Who decides the control flow, the model at run time or a designer ahead of time?
- Does the run need durable, long-running, resumable execution with per-step retries and error handling?
- How much does the job reach beyond the model into other AWS services, fan-out, parallelism, and human-approval waits?
- How provable and auditable does each run need to be?
- Is the GenAI work the whole job, or one step inside a larger process?
The landscape
An agent
The foundation model runs the reason-act-observe loop over an instruction prompt, a set of tools, and optional knowledge bases, choosing each step as it goes. On AWS the platform is Amazon Bedrock AgentCore: a managed harness or your own code on AgentCore Runtime, with tools exposed as MCP tools through AgentCore Gateway. The 2023 service, now renamed Amazon Bedrock Agents Classic, went into maintenance mode on 30 July 2026, closed to accounts with no prior use and frozen on the model catalogue it had then, so new agent work starts on AgentCore rather than on action groups.
It suits a job whose path varies request to request and cannot be drawn in advance. Against an overnight batch what stops it is the missing machinery. There is no declarative per-step retry policy, no fan-out primitive, and no callback token to hold the run open for a two-day human decision, and an AgentCore Runtime session on a microVM lasts at most eight hours. Every one of those becomes yours to build and operate.
The sibling piece on orchestrating multiple agents covers the multi-agent extension of this shape.
A Bedrock Flow
A drawn graph of Bedrock-native nodes, prompts, knowledge bases, agents, Lambdas, conditions, iterators and collectors, wired together with data links. The designer fixes the order and the model works inside a node. The node types, the versioning, and the alias promotion are covered in the deterministic pipeline built as a Flow.
For a single-pass pipeline over Bedrock building blocks that is the least assembly for the most predictability. This job is not that shape. InvokeFlow runs the graph until it finishes or times out at one hour. Asynchronous flow executions, in preview at the time of writing, stretch a run to 24 hours with each node capped at five minutes, but the quotas around them stay tight: 40 nodes in a flow, one iterator node, no per-item retry policy, and no callback that would pause the run for a reviewer.
An AWS Step Functions state machine
The general-purpose durable workflow orchestrator in AWS, and the only one of the three that is not a Bedrock feature. A state machine is a set of states. Task states call a service, a Lambda, or a model. Choice states branch on the data. Parallel states run several branches at once, and wait states hold for a duration or a timestamp. A Map state runs one branch per item in a collection, either inline at up to 40 concurrent iterations or in distributed mode, where each item becomes its own child execution and up to 10,000 run in parallel.
There are two workflow types and the split matters here. A Standard execution is durable, runs for up to a year, is billed per state transition, and keeps an execution history that Step Functions retains for 90 days after the run closes. An Express execution is capped at five minutes, is billed on the number of executions plus their duration and memory, and suits high-volume short-lived work. Express captures no execution history of its own, and it supports neither the callback pattern nor distributed Map.
Failure handling is declared rather than written. A state carries Retry blocks that match error names with their own interval, backoff rate, and attempt limit, and Catch blocks that route a named error to a recovery state. Waiting on something outside the workflow is declared the same way. A task invoked with the wait-for-callback pattern hands out a task token and stays in that state until something calls back with SendTaskSuccess or SendTaskFailure, which turns a two-day human decision into a state rather than a queue you maintain.
Take it when the overnight qualities dominate: surviving failures across hours, fanning out over a collection, reaching across AWS, and waiting for a person, with the model participating rather than conducting. You design and maintain the state machine in return, and for a job that is purely model reasoning that is more scaffolding than the work needs.
Evaluation
Side by side
| Property | Bedrock agent | Bedrock Flow | Step Functions |
|---|---|---|---|
| Control flow decided by | Model, at run time | Designer, ahead of time | Designer, ahead of time |
| Longest single run | 8h AgentCore microVM session | 1h via InvokeFlow, 24h in preview flow executions |
✓ (up to 1 year, Standard) |
| Recovers a failed step for you | ✗ | ✗ | ✓ (Retry and Catch per state) |
| Fan-out across a batch | ✗ | One iterator node per flow | ✓ (Map, distributed to 10,000 children) |
| Pause for days on a human | Build it yourself | ✗ | ✓ (wait-for-callback task token) |
| Reach across AWS services | Tools through AgentCore Gateway | Bedrock-centric plus Lambda | ✓ (over 200 services) |
| Record of what happened | Trace in AgentCore Observability | Node-level execution events | ✓ (execution history, 90 days) |
| Best when | Path must be discovered at run time | Single-pass Bedrock-native pipeline | The model is one step in a durable process |
Reading it for the document job: the sequence is known, so a model-driven engine has nothing to discover; the run is long, retry-heavy, fans out over thousands of records, and pauses for a human, which is more than a Flow is built to carry; the state machine is the fit.
How to route it
The solution
Build it as a Step Functions state machine, with Bedrock as one task among many. Everything the job actually demands, surviving a throttle, surviving a Lambda timeout, waiting two days on a reviewer, and being auditable document by document afterwards, is a first-class primitive there and code you would otherwise write and operate.
Model the document job directly. A Map state in distributed mode fans out across the batch, reading the manifest from S3 and running each document as its own child execution. Inline mode will not do at this size: it caps at 40 concurrent iterations and folds every iteration into the parent’s 25,000-event history, which a few thousand documents would exhaust. For each document, a task state invokes Bedrock to extract the fields, a choice state checks the validation result, a retry policy on the extraction state handles throttling with exponential backoff, and a catch handler routes a repeat failure to a wait-for-callback task that holds on a task token until a reviewer decides.
Two details follow from that shape. The child workflows have to be Standard rather than Express, because the callback pattern is Standard-only. And task input and output cap at 256 KiB, so a long document goes to the model by reference: the optimised bedrock:invokeModel integration accepts Input.S3Uri and writes the response back to Output.S3Uri.
Pick the workflow type on run length. Standard gives durable execution up to a year with a history you can inspect state by state, retained 90 days after the run closes, which is what the operations team needs to open a run and see which documents took which path. Express suits short, high-volume runs and carries no history of its own, so it is the wrong half of the choice here.
You design and maintain the state machine in return. That work is worth doing when the workflow is the product and the model is a participant, and wasted on a job that is purely model reasoning.
Why not an agent. The model works out the branching itself, and that is worth its nondeterminism when the path has to be discovered. This path is known in advance. Asking an agent to carry it means asking it to be a durable workflow engine, and it is not one: building retries, resumability, a two-day human pause, and an audit trail around an agent is rebuilding Step Functions by hand, with none of the guarantees.
Why not a Flow. A Flow draws a fixed graph with little assembly and records events node by node, which suits a single-pass, Bedrock-native pipeline well. This job is none of those things. It runs for hours, fans out across thousands of records with independent per-item retry, and pauses for a human decision. That is the state machine’s territory, not the Flow’s.
Where they combine. The engines layer, and the answer here stays a state machine because of it. A task state calls a model through the optimised bedrock:invokeModel integration and an AgentCore harness through bedrockagentcore:invokeHarness, so a step that later needs the model to choose its own path becomes a harness invocation inside the state machine. The harness integration supports request-response only, and its task state stops at 15 minutes whatever TimeoutSeconds says, so a longer agent turn belongs behind a callback rather than a blocking task.
Worked example
The job: a few thousand supplier documents, extract fields, validate, store, escalate a second failure to a human.
Framed as an agent. One agent per document reads the file, calls an extract tool, checks the schema, and on a second failure calls a tool that notifies a reviewer. It works for a single document in isolation, but the batch has no home. Nothing durably tracks three thousand in-flight runs, nothing retries a throttled model call with backoff on your behalf, and an eight-hour session ceiling rules out a two-day human response. You would wrap the agent in your own queue, retry logic, and state store, which is a workflow engine you are now maintaining.
Framed as a Flow. A Flow expresses the per-document happy path well: input, extract via a prompt or agent node, condition on validity, branch to store or to a notify Lambda. It falls short on the batch shape. One iterator node per flow, a 40-node ceiling, a five-minute cap on each node, and no callback for a reviewer put this batch past what the Flow runtime is meant to carry.
Framed as a state machine. A distributed Map state iterates the batch, one child execution per document. Inside each: a task state invokes Bedrock to extract, with a retry policy for throttling and a catch for hard errors; a choice state branches on the validation outcome; a failed document goes to a wait-for-callback state that holds on a task token until a reviewer resolves it, then rejoins; a clean document writes to the store. The run is durable across the whole night, every document has its own inspectable history, and the model call is one state among many. The other two framings were rebuilding a workflow engine that already exists.
What’s worth remembering
- A sequence you can draw on a whiteboard does not need a model choosing its order, so for a known workflow the real choice is between the two runtimes that walk a drawn sequence.
- An agent runtime has no declarative retry policy, fan-out primitive, or callback pause, and an AgentCore microVM session lasts at most eight hours. Needing those around an agent is a sign the job is really a workflow.
- A distributed Map state runs each item as its own child execution, up to 10,000 at once, and the callback task-token pattern holds a state open for a human decision. Both are Standard-only, and an inline Map caps at 40 iterations sharing the parent’s 25,000-event history.
- The engines combine. A task state can call a model with
bedrock:invokeModelor an AgentCore harness withbedrockagentcore:invokeHarness, so a durable outer workflow can wrap a model-driven inner step. - When the generative work is one step inside a longer business process, that process should own the sequence and the model should be a task inside it. Read that way, the state machine goes on the outside and Bedrock goes in a state.