This is the third lab in the managed track, and the track’s first trip out of text: image generation, video generation, and the asynchronous invocation pattern that video forces on you. The full lab is in lab-13-weekly-creative.zip; unpack it and follow the README.
Before your first lab, do the one-time, once-per-account setup: run the zip’s preflight.sh to confirm your account is ready, then deploy the lab reaper, a standing backstop that auto-deletes any lab you forget to tear down after 24 hours. This lab runs in us-west-2, so point the reaper there too (REAP_REGIONS=us-east-1,us-west-2).
The scenario
Greenbox’s core marketing artefact changes every week, because the box does. What goes in depends on what came out of the ground, so the produce list, the featured farms, the suggested recipes, and the one vegetable subscribers will not recognise are all different by Monday. Somebody has been briefing a designer every Friday, and it is the same brief every time with different nouns in it.
The weekly change already exists as data. Operations publishes a box manifest so the packing sheets, delivery notes, and subscriber emails agree on what is in the box. If the manifest is the source of truth for the box, it can be the source of truth for the picture of the box. This lab wires generation onto the end of that pipeline: four kinds of asset from one JSON file, unattended, in a consistent illustrated house style. Next week’s creative becomes a pull request against a data file.
What you’re given
CloudFormation builds an S3 bucket (the manifest goes in; the stills and video come out) and a Lambda with its role. Neither model appears in the stack, because on-demand generation is serverless: the stack is storage and permissions, nothing else.
The interesting file is manifest.json. Alongside the produce list, the farms, the recipes, and the tricky vegetable, it carries a house_style block, and that block is where two of the lab’s ideas live. The first is that the look is a field rather than a habit. One function assembles a style phrase, a palette phrase, and a register phrase into a tail that every prompt inherits, which is what makes eight different assets read as one family. The second is that the honesty policy is code too. Its negative_text excludes photographs, people, faces, hands, and farmers by name. Greenbox’s rule is that the real growers appear only in real photography. A generated farmer on the page about the farm that grows your carrots would mislead subscribers exactly the way generated walk-through footage of a real house misleads a buyer. Everything this pipeline produces is artwork of produce, recipes, and technique, in a register nobody would mistake for a photograph, and that register carries the whole of the policy on its own.
src/handler.py already loads the manifest, builds every prompt, derives a stable seed per asset, and handles the S3 plumbing. Two gaps are left, and they are the two calls.
Your task
Two gaps in src/handler.py, one per call shape. The stills are the synchronous side: generate_image() is one invoke_model call against IMAGE_MODEL_ID, with a JSON body carrying the assembled prompt, the house negative_prompt, the shared aspect ratio, the derived seed, and PNG as the output format. Read the response body, parse it, and decode the base64 images out of the reply. The module docstring in src/handler.py has the exact request and reply shapes.
Return what arrived, not what you asked for. images, seeds and finish_reasons come back as three lists that line up by position, and null in finish_reasons is the success case. Anything else names what stopped it: Filter reason: prompt, Filter reason: input image, Filter reason: output image, or Inference error. Nothing is raised, so a pipeline that reads images[0] and skips the reasons publishes a frame the filter already rejected. There is no count field either: one call is one image, so a set of assets is a set of calls.
The clips are the asynchronous side, one job each. start_clip() calls start_async_invoke against VIDEO_MODEL_ID: the model input carries the clip’s prompt, the same aspect ratio, the duration and resolution (5s or 9s, 540p or 720p), the loop flag, and a keyframes block whose frame0 wraps the still as base64 with its media type. The output data configuration names the S3 prefix the render should land under, and the invocation ARN in the response is what you hand back. Again, the docstring spells out the exact shape.
The keyframe travels inside the request rather than as a reference to S3, which is why the handler reads the still back out of the bucket itself before starting the job. frame0 is where the clip opens; a frame1 beside it would pin the closing frame too, and leaving it out leaves the five seconds after the opening frame unconstrained. loop is the one setting that changes the shape of the result rather than its content, and the technique clip sets it so a support page can play the same five seconds continuously without a visible jump. The call returns an invocation ARN and nothing else, because a render is a job: get_async_invoke reports Completed, InProgress, or Failed, and the MP4 lands under the prefix you named.
Ray 2 has no multi-shot task, so the website piece is four separate renders rather than one long one. Joining them into a single film is an ffmpeg concat afterwards, and one render per clip means a filtered clip costs one clip to redo rather than the whole film.
Deploy and prove it
Costs split the run in two, which is why the scripts do too. The stills path is cents: deploy.sh, then stills.sh for the hero and the recipe cards, then test.sh, which checks the objects landed and hands you a presigned link to the hero. The motion path is billed per second of output video, so motion.sh prints the clip count, the seconds it implies, and the pricing page, then stops for a typed confirmation before it starts anything. A clip takes a few minutes to render, and the four jobs run at once.
cd lab-13-weekly-creative
./scripts/deploy.sh # stack, manifest, handler. Cents.
./scripts/stills.sh # hero + recipe cards. Cents.
./scripts/test.sh
./scripts/motion.sh # website piece + technique clip. Gated. Real money.
./scripts/teardown.sh
Both models live in us-west-2, which is the lab’s default region for that reason rather than a preference. One prerequisite bites almost everyone: Model access for stability.stable-image-core-v1:1 and luma.ray-v2:0 are two separate console grants, and the failure usually arrives on the second one, after the stills worked and access felt sorted. If you have read about this pipeline running on Amazon Nova Canvas and Nova Reel, it did. Both have been Legacy since 30 March 2026, so an account that was not already using them cannot adopt them at all, and both reach end of life on 30 September 2026. That is a fair preview of the maintenance a generation pipeline needs.
Then test the claim. Change the data: swap the substitution, add a recipe, rewrite the palette. Run deploy.sh and test.sh again and the artwork follows, because the prompts are built by code from manifest fields and each still derives a stable seed from the manifest’s base. A rerun of the same manifest regenerates the same stills; when a picture changes, the data changed. Next week’s creative is a manifest edit, not a design request. This is the structured-output lab inverted: there, prose went in and JSON came out; here, JSON goes in and creative comes out.
When you want the reference answer, deploy it with SRC=solution ./scripts/deploy.sh, or unfold it here:
Show the answer
resp = _bedrock.invoke_model(
modelId=IMAGE_MODEL_ID,
body=json.dumps({
"prompt": prompt,
"negative_prompt": NEGATIVE_TEXT,
"aspect_ratio": "16:9",
"seed": seed,
"output_format": "png",
}),
)
payload = json.loads(resp["body"].read())
images, reasons = payload["images"], payload["finish_reasons"]
video = _bedrock.start_async_invoke(
modelId=VIDEO_MODEL_ID,
modelInput={
"prompt": clip_text,
"aspect_ratio": "16:9",
"duration": "5s",
"resolution": "720p",
"loop": loop,
"keyframes": {"frame0": {
"type": "image",
"source": {"type": "base64",
"media_type": "image/png",
"data": base64.b64encode(keyframe_png).decode()}}},
},
outputDataConfig={"s3OutputDataConfig": {"s3Uri": output_uri}},
)
return video["invocationArn"]
The honesty line, and the keyframe contract
Two rules carry the quality of the result, and neither is a model setting.
The first is the policy in negative_text. Greenbox generates artwork of produce, recipes, and technique, and never people: the real growers appear in real photographs and real footage, because a generated farmer on the website would misrepresent something that exists. Generated media is for illustration sold as illustration, in a register nobody mistakes for a photograph. Bedrock’s watermark detection covers Titan Image Generator G1 and Nova Canvas, and AWS documents no watermark for either model used here, so there is no provenance signal underneath to settle the question later. The register is the only thing holding the line, which is an argument for drawing rather than for a disclaimer nobody reads. The policy lives in the manifest rather than in anyone’s memory, which is what makes it survive the Friday rush.
The second is the keyframe contract: each clip opens on a still and animates away from it, so the still is the only moment of the clip you fully control. The manifest’s motion phrase (“the camera holds almost still”) keeps the generated seconds anchored to the designed one. Prompt a sweeping camera move instead and the output leaves the designed frame behind within a second, which is the walk-through failure again. The technique clip goes one step further and loops, so the last frame has to meet the first: the five seconds close back onto the designed frame rather than only starting from it.
What’s worth remembering
- Stills are synchronous, clips asynchronous. Stable Image Core returns the image from
invoke_model; Ray 2 returns an ARN fromstart_async_invoke, followed withget_async_invoke. - Share one
aspect_ratio. Stable Image Core returns 640 to 1,536 px a side, inside the 512-to-4096 px range a Ray 2 keyframe accepts. - Read
finish_reasonsevery time.nullis success; anything else names the filter or error, and nothing is raised. - Seeds reproduce stills only. A fixed seed gives reproducibility, not consistency; the video request has no seed, so only the opening frame is reproducible.
- Video writes under the caller’s permission. Bedrock delivers to your bucket with the caller’s own
s3:PutObject, unlike a customisation job’s service role. - Hold the house style in code. With no style-preset field, one function assembling medium, palette and register keeps eight prompts consistent.