The situation
A team has fine-tuned Meta Llama 3.1 8B Instruct on Bedrock against their own support transcripts, in us-west-2, the one Region where that customisation runs. It tests well: noticeably better at their domain than the base model with a long prompt. They are ready to put it behind a live feature. The first surprise is that the on-demand invoke path they used all through prototyping, pay per token with no capacity to manage, does not serve this model. Bedrock deploys some custom models for on-demand inference, but a fine-tuned Llama 3.1 8B is not one of them, so Provisioned Throughput is the only way to serve it.
The traffic is not flat. Weekday business hours carry the bulk of it, with a sharp mid-morning peak when the support queue fills, near silence overnight, and a long quiet tail at weekends. Someone has pulled a number for the busiest minute: roughly the token volume the feature has to sustain when the queue is at its worst. Finance wants the cheapest per-unit rate, which means a six-month commitment. Engineering has been caught before by locking in capacity a fortnight before a traffic pattern changed, and asks what the commitment covers and what it forecloses.
Underneath the calendar question is a sizing question. Reserve too little and the mid-morning peak throttles real users; reserve too much and idle units bill around the clock for throughput nobody consumes. An account also starts with no model units at all, so the whole exercise opens with a support request for the units themselves.
What actually matters
For this model it is not a provisioned-versus-on-demand choice, and whether it is one for yours depends on the base model you customised. A custom model deployment serves a fine-tuned model on demand, billed per token, for models customised from Nova Micro, Nova Lite, Nova Pro or Nova 2 Lite in us-east-1, and from Llama 3.3 70B Instruct in us-west-2, where the customisation job ran on or after 16 July 2025. Llama 3.1 8B is on neither list, so reserved capacity is the only path, priced per unit the same as the base model it came from. Import your own weights and the bill is different again: Bedrock charges custom model units per minute, over five-minute windows, and runs more or fewer copies of the model as demand moves. Three cost shapes, decided by which model you started from.
Capacity is reserved in model units, and the unit is the thing to understand properly. One unit delivers a defined throughput for one specific model: the input tokens it can process across all requests in a minute, and the output tokens it can generate in that minute. It is not a share of a pool, and it is not burstable; it is a fixed rate you have reserved. Two units give you twice the rate. AWS does not publish the per-unit token figures, so the sizing arithmetic starts with the numbers your AWS account team gives you for your model.
Then the cost shape, which is the trap. You are billed hourly for reserved units, whether or not traffic fills them. A unit sized for a peak that lasts ninety minutes a day still bills for the other twenty-two and a half hours. That is the over-provisioning failure: capacity sized to the worst minute, billed around the clock, mostly idle. The opposite failure is under-provisioning, where the reserved rate sits below the real peak and requests over it are throttled, so the mid-morning surge turns into errors and retries for actual users. Sizing sits between those two failures, and headroom is the guard against the second.
The unit count is fixed once the reservation exists. The update API changes the name of a Provisioned Throughput and, for a custom model, which model it points at. It does not change the number of units. Resizing means purchasing a second reservation at the new count and deleting the first, which a commitment blocks until the term is over.
The term is a separate lever from the unit count, and it trades rate against flexibility. Three levels are available, and custom models support all three: no commitment, billed hourly at the highest per-unit rate and deletable at any time; a one-month commitment at a lower rate; a six-month commitment at the lowest rate. A commitment cannot be deleted before its term ends, and a reservation renews automatically at the end of each term. Leaving it alone at the boundary is a decision to take another term.
The last thing that matters is that demand shape decides how well any of this fits. Steady, predictable load maps cleanly onto reserved units and suits a commitment, because the units you are billed for are the units you use. Spiky load with deep troughs is the awkward case: size to the peak and the troughs bill for nothing, size to the average and the peaks throttle. Neither the unit count nor the term fixes a genuinely spiky profile on its own; it just moves where the pain sits.
What we’ll filter on
- Serving eligibility: does this base model have a custom model deployment path, or is Provisioned Throughput the only way to serve it?
- Peak throughput: what is the busiest-minute demand in input and output tokens per minute, and what does one model unit deliver for this model?
- Demand shape: steady and predictable, or spiky with long idle troughs?
- Commitment appetite: how confident is the traffic forecast over one month, and over six?
- Cost of idle versus cost of throttling: which failure hurts this feature more, a bigger bill or dropped peak requests?
The landscape
On-demand invocation. Pay per token processed, no capacity to reserve, no floor, no commitment. This is the default for base foundation models and it is where the prototype lived. Bedrock extends it to custom models through a custom model deployment, which you invoke by the deployment ARN, but only for the Nova text models and Llama 3.3 70B Instruct. A fine-tuned Llama 3.1 8B drops out of contention here regardless of how attractive the billing shape is.
Provisioned Throughput, no commitment. Reserve model units billed by the hour, at the highest per-unit rate, and delete the reservation as soon as the numbers justify it. Resizing is a replacement rather than an edit: purchase the new unit count, move traffic to the new provisioned model ARN, delete the old reservation. This is the term for a workload whose shape you do not yet trust, or one you expect to run only for a bounded window.
Provisioned Throughput, one-month commitment. The same reserved units at a lower per-unit rate in exchange for holding them for a month. A reasonable middle when the near-term traffic is understood but the half-year is not, and a common way to run a steady production workload without a six-month lock on it.
Provisioned Throughput, six-month commitment. The lowest per-unit rate, locked for the term. This is the right call only when demand is both steady and confidently forecast that far out, because the reservation can be neither resized nor deleted inside the term, and it renews for another six months unless you delete it at the boundary.
The unit count is orthogonal to the term. Each option above is purchased as some number of model units, and the number comes from the same peak-tokens-per-minute arithmetic in every case. The term sets the rate and the lock; the unit count sets the ceiling.
Evaluation
Side by side
| Option | Serves this custom model | Per-unit price | Earliest exit | Idle-cost exposure | Best for |
|---|---|---|---|---|---|
| On-demand deployment | ✗ | Per token, no floor | Nothing reserved | None | Nova text models, Llama 3.3 70B |
| PT, no commitment | ✓ | Highest | Delete at any time | Per reserved unit-hour | Unproven traffic, bounded-window runs |
| PT, one-month | ✓ | Lower | End of the month | Per reserved unit-hour | Understood near-term production load |
| PT, six-month | ✓ | Lowest | End of six months | Per reserved unit-hour | Steady, confidently forecast load |
Reading the table against the situation: on-demand is unavailable for a model fine-tuned from Llama 3.1 8B, so the whole decision is which Provisioned Throughput term to take and how many units to reserve. Finance is reaching for the bottom row; engineering’s caution about the forecast is an argument for one of the middle two until the shape is proven.
The solution
Start with the unit count, because the term does not matter if the ceiling is wrong. Take the busiest-minute demand in input and output tokens per minute, divide each by what one model unit delivers for this model, and use whichever side needs more units. That gives the raw coverage for the peak. Then add headroom rather than rounding to the measured peak. The peak you measured is an average over a minute, real traffic is burstier inside that minute, and a fine-tuned model’s output length drifts as prompts evolve. Headroom protects against throttling, and the margin is a judgement about how spiky the minute really is and what a throttled request does to the feature. Size to the peak plus that margin, not to the daily average, or the mid-morning surge throttles every day.
Now the idle problem the peak sizing creates. A unit count set to the worst minute bills at that level for all twenty-four hours, including the overnight silence and the weekend tail. Nothing follows the curve down for you: reserved units stay reserved until you delete the reservation, and the count cannot be changed while it exists. If the trough is deep and long, the honest question is whether throttling at a lower unit count is genuinely worse than idle billing at the peak count. Sometimes accepting a little throttling at the very tip of the peak, and sizing below the absolute maximum, costs less overall than billing all night for headroom used ninety minutes a day. That is a per-feature call, and it turns on whether a throttled request degrades gracefully with a retry or hard-fails a user.
The term is the last decision and the reversible-versus-locked one. If the traffic forecast is honest only a few weeks out, no commitment or one month keeps the reservation short-lived while you watch the real curve, at a higher hourly rate per unit. Once a month or two of production data shows the peak is stable, a longer commitment on that proven capacity gets the lower rate. A defensible pattern is to hold the steady floor on a longer commitment and the uncertain margin on a separate, shorter reservation, so the lock only ever covers demand you are confident in. Jumping straight to six months on day one, before any production traffic has been seen, is the move most likely to end in idle units you cannot delete, or a lock that no longer fits the curve. Diary the renewal date as well, because a commitment takes another term on its own.
None of that sizing survives contact with real traffic unless you measure it, and a reservation gives you an obvious ceiling to measure against. Graph InputTokenCount and OutputTokenCount for the provisioned model against the reserved tokens per minute on each side, and the ratio between them is your utilisation. Put InvocationThrottles next to it, because utilisation approaching the ceiling and utilisation actually hitting it are different states, and only the throttle counter tells you which one the mid-morning peak is in. Hold all three as percentiles over a fortnight rather than averages, since a reservation sized against a mean is wrong at the peak by construction. The daily mean of a workload that is silent overnight describes nothing anybody experiences. This is a row on the same CloudWatch dashboard that watches a Bedrock feature’s cost and latency rather than new instrumentation.
The two readings point at different fixes. Sustained utilisation well under the reservation across the whole day, with the throttle counter flat, is over-provisioning: you are billed for a ceiling nobody reaches, and the remedy is a smaller reservation, which means a replacement at the first boundary the term allows. Throttles at the peak alongside low utilisation for the rest of the day is a shape problem rather than a size problem. Another model unit would fix it while billing around the clock for capacity used ninety minutes a day, so moving batch work off the reservation, so it stops competing with the interactive peak, is usually the cheaper answer. Capacity planning for token processing is that judgement: working out which of the two the numbers describe before reaching for more units. Prompt and completion patterns also drift as the feature evolves, longer contexts in and longer replies out, so utilisation review belongs on the commitment boundary rather than once before launch. Arriving at each renewal with a fortnight of percentiles instead of a hunch is most of what optimising Provisioned Throughput amounts to.
One base-model note, because the two cases are easy to blur: a base foundation model can also be put on Provisioned Throughput, usually to guarantee a throughput floor for a latency-sensitive or high-volume workload that on-demand quotas would throttle. That is a legitimate but uncommon choice, since most base-model traffic is better served on demand. The asymmetry is what to hold onto: a base model may use Provisioned Throughput, and a customised model must, unless its base model is one of the few with a deployment path.
Worked example
Call the busiest minute the moment the support queue peaks mid-morning. Suppose measurement puts that minute at a demand the team can express as tokens per minute in and out, and the per-unit throughput their account team quoted covers a known fraction of it, so the peak divides out to three units of raw coverage. Output tokens turn out to be the binding side, because the drafted replies are long, so the arithmetic is done against the output rate.
Rounding to three units exactly would meet the average of the peak minute and throttle the bursts inside it, so they size to four: three for the measured peak, one for headroom against intra-minute spikes and output drift. Four units it is, as the ceiling.
Then the idle question. Those four units bill all night and all weekend, when demand is near zero. The team looks at the curve and decides the deep trough does not justify a second, smaller off-peak reservation, because deleting and repurchasing capacity twice a day costs more in operational effort than it saves. They note it as a lever if the bill grows.
On the term, they hold back from six months. The feature is new and the peak could move as adoption grows, so they take the four units on a one-month commitment: a lower rate than no commitment, and only a month of lock. Two months of production data later, three of those four units are demonstrably the stable floor and the fourth is genuine swing capacity. At the next boundary they delete the four-unit reservation and purchase two: three units on a six-month commitment at the lowest rate, and one unit with no commitment, covering the part of the curve they are least sure of. The lock only ever covers demand they have actually seen.
What’s worth remembering
- Check the base model before planning capacity: a custom model deployment serves Nova text models and Llama 3.3 70B Instruct on demand per token, and a model fine-tuned from anything else on Bedrock is served through Provisioned Throughput.
- Reserved units are billed by the hour they exist, filled or idle, so peak-sized capacity costs the same overnight and at weekends as it does mid-morning.
- Size the unit count from busiest-minute tokens per minute, on whichever of input or output needs more units, then add headroom; sizing to the daily average throttles the peak.
- A reservation’s unit count is fixed at purchase, so resizing means buying a new Provisioned Throughput and deleting the old one, which a commitment prevents until the term ends.
- The terms are no commitment, one month and six months, cheapest per unit in that order, and a reservation renews automatically unless you delete it at the boundary.