Exam Room · Advanced Generative AI Developer

Cheat Sheet: Model Selection and Inference

· 22 min read

Generative AI Development · part of The Exam Room

A fast revision sheet for model choice and inference options across Amazon Bedrock and self-hosted SageMaker. Model lifecycle dates move, so check the model card before you commit.

Services at a glance

Thing What it is Reach for it when
Amazon Nova Micro Text-only. 128K context, 10K max output. Lowest cost and latency in the family. High-volume text work where speed and price decide it.
Amazon Nova Lite / Pro Multimodal in (text, image, video), text out. 300K context, 10K max output. Images or video in the input. Pro for accuracy, Lite for cost.
Amazon Nova 2 Lite The current Nova generation. Up to 1M context and 65,536 output tokens, extended thinking, built-in web grounding and code interpreter. New multimodal work, agentic flows, long documents.
Amazon Nova Multimodal Embeddings One embedding space over text, images, documents, video and audio. Cross-modal retrieval and multimodal search.
Amazon Nova Canvas / Nova Reel Image and video generation. Both Legacy, both withdrawn on 30 September 2026. Nothing new. Check the catalogue before you promise a generator.
Amazon Titan Embeddings. Titan Text Embeddings V2 takes 8K tokens and has configurable output dimensions. Text embeddings for retrieval and search.
Anthropic Claude Strong general reasoning, long context, tool use. Complex reasoning, tool use, long documents.
Meta Llama Open-weight text models, plus multimodal Llama 4 Scout and Maverick. An open-weight preference inside managed Bedrock.
Mistral Efficient text models and larger reasoning tiers. Cost-efficient text; European provider preference.
Cohere Embed v4, Embed English and Multilingual, Rerank 3.5. Embeddings and reranking for retrieval.
AI21 Jamba Long-context text. Jamba 1.5 Large and Mini are Legacy, EOL 26 November 2026. Nothing new.
Stability AI The Stable Image family: inpaint, outpaint, erase, upscale, style and structure control. Editing, or generation steered by a sketch, structure or style reference.
The wider catalogue DeepSeek, Qwen, OpenAI gpt-oss, Google Gemma, Writer, xAI, MiniMax, Moonshot, NVIDIA, Z.AI, TwelveLabs. A named third-party or open-weight model without leaving Bedrock.
On-demand inference Pay per token, no commitment, quota-bound. Spiky or unpredictable traffic; getting started.
Service tiers Standard, Priority, Flex and Reserved, set per request with service_tier. Sorting latency-sensitive traffic from work that can wait.
Provisioned Throughput Model units billed hourly, with no commitment or a one or six month term. Steady high volume, and most customised models.
Batch inference Asynchronous S3 to S3, at half the on-demand token price. Large offline jobs with no latency requirement.
Cross-Region inference profile Routes a request across Regions, scoped to a geography or global. Bursty load above one Region’s quota.
SageMaker hosting Self-host open or custom models on your own endpoints. Models not on Bedrock, or full control of serving.
Intelligent Prompt Routing One endpoint over two models in the same family, routed on predicted response quality. Mixed request difficulty behind a single endpoint.
AWS AppConfig Dynamic configuration a running service reads at request time, deployed with a bake time and CloudWatch-alarm rollback. Holding the model id, inference configuration and prompt version outside the code, so a swap is a configuration deployment.
Well-Architected Generative AI Lens A custom lens you import into the Well-Architected Tool and review a workload against. A tracked improvement plan with dated milestones, rather than a one-off opinion.

Decision rules

  • Text-only and high-volume: start at Nova Micro. Move up only when quality falls short.
  • Images or video in the input: use a multimodal model, such as Nova 2 Lite, Nova Lite or Pro, Claude, or Llama 4. Micro and the embedding models do not accept those inputs.
  • Generating images or video: read the current catalogue first. Nova Canvas and Nova Reel go on 30 September 2026, and Stability AI’s Bedrock line-up is editing and guided generation rather than a plain text-to-image base model.
  • Embeddings for retrieval: Titan Text Embeddings V2, Cohere Embed v4, or Nova Multimodal Embeddings when the corpus spans modalities. Not a chat model.
  • Narrowing a shortlist: published benchmarks rank models against someone else’s tasks, so use them to cut the field. A Bedrock evaluation job over your own Golden datasetA versioned set of representative inputs with known-good expected outputs, run on every prompt or model change to catch regressions. settles which one ships.
  • Before any score is compared, run the limitation evaluation. Drop the models that cannot do the job at all: context window, maximum output tokens, modalities, tool use, languages, Region and cross-Region availability, and whether fine-tuning and Provisioned Throughput are supported.
  • Spiky traffic and low commitment: on-demand.
  • Within on-demand, four service tiers share one API. Standard is the default, Priority costs a premium for the fastest responses, Flex takes a 50% discount for work that tolerates longer processing, and Reserved holds tokens-per-minute capacity for a one or three month term. Priority, Standard and Flex draw on the same quota; Reserved capacity is separate.
  • Steady high volume: weigh the hourly model-unit rate for Provisioned ThroughputReserved Bedrock capacity bought by the hour for a fixed term, paid for whether traffic fills it or not. against on-demand token spend. A one or six month term lowers the hourly rate.
  • Fine-tuned on Bedrock: Nova Micro, Lite, Pro and Nova 2 Lite deploy on demand in us-east-1, and Llama 3.3 70B Instruct in us-west-2, if the model was customised on or after 16 July 2025. Anything else runs on Provisioned Throughput.
  • An imported model is billed differently again: Custom Model Units per model copy, over five-minute windows from the first successful call, with copies scaled to demand.
  • Large offline job with no latency need: Batch inferenceSubmitting a bulk job of model calls to run asynchronously at a lower per-token price, trading immediacy for cost., at half the price.
  • One Region’s throughput quota is the limit: enable a Cross-region inferenceLetting a request be served from any of several regions, raising effective throughput and riding out pressure in one of them. profile. Global routing runs about 10% cheaper than a geographic profile.
  • Data must stay in one geography: choose the geographic profile over the global one, and check the destination Regions.
  • The model is not on Bedrock: host it on SageMaker.
  • A SageMaker endpoint idle between bursts: serverless, which scales to zero and carries cold starts.
  • Large payloads or slow processing: SageMaker asynchronous inference, which queues requests up to 1GB and an hour.
  • Scoring a whole dataset once with no live endpoint: SageMaker batch transform.
  • Many models on shared infrastructure: inference components for per-model resources and scaling, or a multi-model endpoint for a long tail of similar models sharing one container.
  • Requests vary in difficulty: put Intelligent Prompt Routing in front of two models in one family. It predicts response quality per request and routes on the criteria you set, with a fallback model as the anchor.
  • Failures transient and sparse: retry with exponential backoff and jitter. If the dependency is failing most calls, open a circuit breaker and shed load until a probe succeeds.
  • Circuit-breaker state belongs in a shared store such as DynamoDB, so every Lambda invocation and Step Functions execution reads the same state. State in process memory resets on the next cold start, and each concurrent worker trips independently.

Traps

  • Nova Micro is text-only. Picking it for an image or video task is wrong.
  • Legacy is a hard state. New customers cannot adopt the model, an existing account can lose access after 15 days of inactivity, and no new Provisioned Throughput can be created for it.
  • Lifecycle dates cluster. Nova Premier and the original Nova Sonic reach EOL on 14 September 2026, Nova Canvas and Nova Reel on 30 September 2026, and Jamba 1.5 on 26 November 2026.
  • A sheet written from memory still lists models that are gone. Titan Image Generator G1 v2 passed EOL on 30 June 2026, and Cohere Command R and R+ on 19 August 2026.
  • Video generation is asynchronous. Nova Reel runs through StartAsyncInvoke and writes to S3, so there is no synchronous response to wait on.
  • Provisioned Throughput does more than lower the unit cost. Outside the five base models that support on-demand custom deployment, it is the only way to serve a customised model.
  • Batch inference is asynchronous and offline. It does not lower latency for live requests, it is unavailable for provisioned and imported models, and it supports neither tool calling nor structured output.
  • Cross-Region inference profiles move data between Regions, which can breach a residency requirement. They also cannot be combined with Provisioned Throughput.
  • On-demand is quota-bound. Throttling means raising the quota, reserving capacity, or moving to Provisioned Throughput, not retrying blindly.
  • SageMaker serverless has cold starts. Do not pick it when consistent low latency matters.
  • Model access is now enabled by default in commercial Regions. The real prerequisites are AWS Marketplace permissions on the first call, a valid payment method, and Anthropic’s first-time-use form; a gap in any of them returns AccessDeniedException. GovCloud still enables models by hand.
  • Model availability differs by Region. A model in one Region may be absent in another.
  • Bigger is not automatically better. The smallest model that clears the quality bar is usually the right pick on cost and latency.
  • SageMaker batch transform has no persistent endpoint. Do not use it for real-time serving.
  • Asynchronous inference is for large payloads and long jobs, not a substitute for provisioned real-time throughput.
  • A hardcoded model id looks like a one-line change. Once eight services call Bedrock it is eight code edits, eight reviews, eight releases, and a window where half the fleet is on each model. Held in AppConfig, the swap is one configuration deployment with a bake time and an automatic rollback.
  • Flattening per-model inference parameters into one shared blob breaks the models that do not share them. Names and ranges differ by provider, so an indirection layer needs an inference configuration per model.

Say it in one line

  1. Nova Micro is text-only, Lite and Pro are multimodal, and Nova 2 Lite carries 1M of context with 65,536 output tokens.
  2. Nova Canvas and Nova Reel leave on 30 September 2026, so confirm what generates images or video before you design around it.
  3. Titan, Cohere and Nova Multimodal Embeddings give you embeddings; use them for retrieval and search, not chat.
  4. On-demand charges per token with no commitment and is bound by service quotas.
  5. Standard, Priority, Flex and Reserved are one parameter on the same call, and Flex is half the Standard price.
  6. Provisioned Throughput reserves Model unitThe billing block Provisioned Throughput is sold in – one unit delivers a fixed tokens-per-minute rate for a specific model. per hour and serves every customised model that cannot be deployed on demand.
  7. Custom model deployment covers Nova Micro, Lite, Pro and Nova 2 Lite in us-east-1, and Llama 3.3 70B in us-west-2.
  8. Batch inference is asynchronous S3 to S3, half price, offline, and without tool calling.
  9. Cross-Region profiles raise throughput; geographic ones hold the geography, global ones save about 10%.
  10. Marketplace permissions, a payment method and regional availability are the checks before a first call.
  11. SageMaker real-time serves low-latency traffic on an always-on endpoint, and serverless trades cold starts for scaling to zero.
  12. SageMaker asynchronous inference queues payloads up to 1GB for up to an hour; batch transform scores a dataset with no endpoint at all.
  13. Inference components give each model its own CPU, accelerators, memory and copy count, down to zero copies; multi-model endpoints share one container and load models on demand, with a cold start on the rare ones.
  14. Pick the smallest model that clears the quality bar, and move up only when it does not.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.