Exam Room · Advanced Generative AI Developer

Cost Attribution and Tagging for GenAI Workloads

· 39 min read

Generative AI Development · part of The Exam Room

The situation

A platform team runs a shared Amazon Bedrock estate for the whole company. Three product teams call the same Claude and Nova models through one set of AWS credentials: a support-summariser, a marketing copy generator, and a multi-tenant chat feature that serves a few hundred paying customers. Titan text generation, which two of them started on, is no longer offered; the Titan name survives on Bedrock only for embeddings and image models. The support-summariser has since moved off the shared Claude endpoint onto weights the team fine-tuned themselves and imported into Bedrock. Around the model calls sit the usual scaffolding, Lambda functions for orchestration, an S3 bucket of source documents, and a Bedrock Knowledge Base backing the chat feature’s retrieval.

The monthly bill has grown past the point where anyone waves it through. Finance can see the total Bedrock spend, and that it is up forty per cent quarter on quarter. What they cannot see is which of the three products drove the rise, or which chat customers are heavy enough to be unprofitable. On-demand Bedrock usage lands in the bill as one undifferentiated line for input and output tokens per model; there is nothing in that line that says “marketing” or “tenant 412”. The Provisioned ThroughputReserved Bedrock capacity bought by the hour for a fixed term, paid for whether traffic fills it or not. commitment the chat team bought last quarter is another flat line of hours, and the imported summariser bills on a third clock again. Those two at least belong to an identifiable owner, since each is a thing that exists in the account; neither shows which tenant used the capacity.

The ask is concrete. Attribute the spend per product for internal chargeback, and the chat spend per tenant so the pricing team can find the loss-makers. Alert each team before it blows through its monthly allowance, rather than a fortnight after. Routing each product into its own AWS account just to read the bill is off the table.

What actually matters

The first thing to be clear about is where Bedrock cost comes from, because you can only attribute what you can measure. On-demand inference is priced per thousand tokens, and input and output tokens are priced separately, with output usually the dearer of the two. A verbose summary is not the same cost as a terse one even for identical input. On top of that, any Provisioned Throughput you have bought is charged by the Model unitThe billing block Provisioned Throughput is sold in – one unit delivers a fixed tokens-per-minute rate for a specific model. hour whether or not you send it traffic. That fixed cost has to be spread across whoever the commitment was for. A model whose weights you supplied yourself is on a third clock. Each live copy bills per custom model unit per minute, in five-minute windows from the first successful call, rather than by the tokens it produced, with a monthly storage charge per unit on top. Batch inferenceSubmitting a bulk job of model calls to run asynchronously at a lower per-token price, trading immediacy for cost. runs at half the on-demand rate for work that can wait. All of it aggregates per model unless you give AWS a reason to split it.

That reason is the second thing that matters: the grain you attribute at. Account-level attribution is the crudest, one bill per account. Application-level attribution answers “which product”, and it is the grain chargeback usually needs. Tenant-level attribution answers “which customer”, which is what a multi-tenant SaaS needs to price fairly and spot the unprofitable accounts. Request-level attribution answers “exactly which call cost what”, the grain for anomaly hunting and reconciling a disputed number. Each finer grain costs more to capture and store. Pick the coarsest one that answers the question.

The third is where the signal comes from, because Bedrock has several attribution mechanisms and they answer different questions. Cost allocation tags flow tag keys from your resources into the billing pipeline. Anything taggable, Lambda, S3, a Knowledge Base, can carry a team or cost-center tag and show up split that way in Cost Explorer and the Cost and Usage Report. But a raw on-demand InvokeModel call against a shared foundation model is not a tagged resource, so tags alone cannot split that spend. The asymmetry to hold on to is that some ways of serving a model create a thing in your account, a reserved block of capacity or weights you brought yourself, and such a thing is taggable like any other. Calling a model AWS hosts for everyone creates nothing. Two mechanisms close that gap. An application inference profile is a Bedrock resource referencing one foundation model; you tag it, route a product’s calls through it, and the usage attaches to the profile’s tags. IAM principal attribution needs no resource at all, because Bedrock records the calling user or role on every request, and tags on that identity reach Cost Explorer and CUR 2.0 once activated. Either way one Bedrock line becomes per-application or per-identity lines with no separate accounts.

The fourth is granularity of the raw record. Cost Explorer and the CUR break cost down by tag, service, and time, the billing-grade view finance reconciles against. Neither gives the token count of an individual request. For per-request chargeback, or attributing cost inside one profile to an end user, you need model invocation logging. It writes each call’s token counts and the caller’s identity ARN, and optionally the payloads, to CloudWatch Logs or S3. Request metadata rides in the same record: up to sixteen key-value pairs the caller sets per call, in the X-Amzn-Bedrock-Request-Metadata header or the Converse requestMetadata field. That is the fine-grained ledger you compute a per-request or per-user cost from, and you own the log storage and the arithmetic.

The last thing that matters is closing the loop, because attribution nobody acts on is just a nicer-looking bill. Once spend is split by tag, AWS Budgets can watch each product’s or tenant’s slice and fire an alert, or an automated action, at a threshold. It also forecasts against the trend, so the warning arrives before the month closes rather than after. The reduction levers, a cheaper model for the easy calls, prompt and response caching, batch inference for the non-urgent work, only become targetable once you know where to point them.

What we’ll filter on

  1. Attribution grain, do we need account, per-application, per-tenant, or per-request?
  2. Signal source, does the mechanism reach shared on-demand inference, or only resources that exist in the account?
  3. Billing-grade versus computed, does it reconcile against the AWS bill, or is it a number we derive ourselves from logs?
  4. Setup and running cost, tag hygiene, a resource per model, log storage and processing.
  5. Closes the loop, can it drive an alert or an automated action before the bill lands?

The landscape

Separate AWS accounts. The bluntest split: give each product its own account and let Organizations consolidate the billing. Attribution per product follows automatically, and blast radius and quotas are isolated too. But it does nothing for per-tenant attribution inside a multi-tenant product, and retrofitting it onto a shared estate is a migration, not a config change. Right for hard isolation, overkill purely to read a bill.

Cost allocation tags. Activate tag keys in the Billing console and they become dimensions in Cost Explorer and the CUR. Tag the Lambda functions, the S3 buckets, and the Knowledge Base with team and cost-center, and the cost of that scaffolding splits cleanly per product. The gap is shared foundation-model inference: a plain on-demand InvokeModel is not a resource you can hang a tag on, so tags alone leave that token spend, usually the biggest number, unsplit. The exception matters. A provisioned throughput and an imported model are both resources, each with its own ARN and tags set at creation, so their spend splits by tag with no further machinery. Essential for the surrounding resources, sufficient for capacity you own, insufficient for shared on-demand inference.

Application inference profiles. A Bedrock resource that references one foundation model, or a system-defined cross-Region profile over one, and carries its own ARN and tags. Route a product’s invocations through their profile and the token usage and cost attach to it. Activate the profile’s tags as cost allocation tags and per-application Bedrock spend appears in Cost Explorer and both CUR versions, aggregated per usage type per day rather than per call. Two constraints shape how far it goes. A profile references exactly one model, so the resource count is one per model per attribution unit, and grows again with every model version. And profiles serve InvokeModel and Converse only: a Responses or Chat Completions call naming one is rejected with a 400, and an imported model cannot sit behind a profile at all, so a product on your own weights needs another route. For the OpenAI-compatible APIs on the bedrock-mantle endpoint, a Bedrock project carries the tags instead, set once on the client and spanning every model it calls.

IAM principal attribution. Bedrock records the calling IAM user or role on every inference request, on both endpoints, with no code change and no resource to create. Tags on the role, or STS session tags passed at AssumeRole, become cost allocation tags under the iamPrincipal/ prefix once activated, and a CUR 2.0 export configured to include the caller identity carries the ARN itself. For a gateway fronting many tenants, assuming the role once per tenant with the tenant as a session tag, then caching those credentials, gives per-tenant spend with no resource each. The grain is again per usage type per day, and session tags reach billing only, never the invocation logs.

Model invocation logging. Turn it on per Region and Bedrock writes every invocation’s record, including input and output token counts and the caller’s identity ARN, to CloudWatch Logs or an S3 bucket, with the request and response payloads optional. Request metadata travels in the same record: up to sixteen key-value pairs the caller sets on each call, so one shared client can stamp a tenant or campaign without a resource per value. This is the only source of a genuine per-request token count, so it is what you compute fine-grained chargeback or a per-end-user cost from. It is a computed number, not a billing-grade one; you multiply logged tokens by the published price yourself, and you own the log storage and the query cost. Right for per-request and per-user granularity, more than you need if per-product is the whole question.

AWS Cost Explorer and the Cost and Usage Report. The reporting surface over everything the tags and profiles feed. Cost Explorer is the interactive one, filtering and grouping by the team tag, by service, by month, and it answers most of the attribution questions by itself. The CUR is the export path underneath it, an exhaustive line-item file, hourly or daily, landing in S3 for Athena or Amazon Quick Sight. Reach for it when a custom join is needed, such as Bedrock cost against your own tenant table. Both are billing-grade; neither carries per-request token detail.

AWS Cost Anomaly Detection. A monitor built on the same cost allocation tag dimension the profiles and tags have just established. A budget fires when a number crosses a line somebody chose in advance. An anomaly monitor fires when a slice diverges from its own learned baseline, which catches a doubling inside a small tenant that is invisible against the account total. It runs about three times a day over Cost Explorer data, so an anomaly can take up to a day to surface, and a new monitor needs a day before it detects anything. Right for the “something changed and nobody told us” case, no help at all for enforcing an allowance.

AWS Budgets. A cost or usage budget scoped by the same tags, with alert thresholds and forecasting. This is the loop-closer: a budget per product tag that emails and pages at eighty per cent of the monthly allowance, and forecasts an overrun before it happens. A budget action applies a deny IAM policy or a service control policy at a threshold, either automatically or after someone approves it. It does not attribute anything itself; it watches the slices the tags and profiles have already carved out.

Evaluation

Side by side

Mechanism Finest grain Splits shared on-demand inference Billing-grade Per-request tokens Closes the loop
Separate accounts Per product ✓ (one bill each) ✓ ✗ ✗
Cost allocation tags Per resource ✗ shared, ✓ capacity you own ✓ ✗ ✗
Application inference profiles Per app, one model each ✓ foundation models only ✓ ✗ ✗
IAM principal attribution Per identity or session tag ✓ both endpoints ✓ ✗ ✗
Model invocation logging Per request ✓ (computed) ✗ ✓ ✗
Cost Explorer / CUR Per tag, per hour reports it ✓ ✗ ✗
AWS Cost Anomaly Detection Per tag slice reports it ✓ ✗ ✓
AWS Budgets Per tag reports it ✓ ✗ ✓

Reading the table against the three asks. Per-product chargeback wants application inference profiles for the products on shared models, plus cost allocation tags on the scaffolding and on any capacity the teams own outright, all surfaced in Cost Explorer. Per-tenant profitability is better served by session tags on the identity than by a profile each, and by model invocation logging where the grain has to be finer than a day. Per-request or disputed numbers require invocation logging. Every one of those slices should carry a Budget so the alert beats the bill, with an anomaly monitor on the same dimension for the spikes nobody thought to set a threshold for. No single mechanism does the whole job.

The solution

The per-product chargeback is the application inference profile case, the one that finally splits the token spend. Create a profile for each product and model pair, since a profile references exactly one model, and tag each with team and cost-center. Change each product’s Bedrock client to invoke via the profile ARN instead of the bare model ID. Activate those tag keys as cost allocation tags in the Billing console, and within about a day Cost Explorer starts showing Bedrock cost grouped by team. Activation is not retroactive, so only spend after that moment is tagged. In the same pass, tag the Lambda, S3, and Knowledge Base resources with the matching keys so the scaffolding cost lands in the same buckets. That covers the two products calling shared models; the summariser on imported weights takes the route below. Tag governance is what makes it work. A call routed through the wrong profile, or a resource left untagged, shows up as unattributed spend. Enforce the tags with a Service Control Policy or a tag policy.

The summariser is simpler than that, because its model is a resource the team owns. Set team and cost-center as tags on the imported model when the import job runs, and activate the keys. The per-minute charge for its live copies and the monthly storage charge then land against that team in Cost Explorer, with no profile and no routing change in the application. Treat the import job as the moment attribution is decided rather than something to retrofit. The same holds for the chat team’s provisioned throughput: tag the reserved capacity and its hours attribute to the team that committed to it. What neither can do is split spend inside itself, since one resource carries one tag however many products or tenants call it.

The per-tenant profitability question is better answered through identity than through more profiles. Have the chat feature’s gateway assume its Bedrock role once per tenant, passing the tenant as an STS session tag and caching the credentials for the session. Activate that tag key, filtered by type IAM principal, and tenant spend arrives in Cost Explorer and CUR 2.0 at billing grade with no per-tenant resource at all. The trust policy has to allow sts:TagSession for the tag to flow.

Where the grain has to be finer than a day, or the product runs on imported weights, model invocation logging is the fallback. Set the tenant ID in request metadata on each call, log every invocation’s token counts, and compute per-tenant cost by joining the logs to the published per-token prices in Athena. That is a derived number, not a billing-grade one, so treat it as the management view for pricing decisions rather than the figure finance reconciles against. Nothing in Bedrock enforces request metadata, so set it in a shared client rather than trusting each caller. Logging the payloads as well as the counts brings tenants’ prompt content into your logs, which is a data-handling decision to make deliberately.

Logging is also the answer whenever the question is per-request: which call spiked the bill on the third, or how to reconcile a tenant’s disputed invoice. Cost Explorer and the CUR stop at the tag and the hour. Log volume at scale is its own bill, so scope logging to the Regions and models that matter, and expire the logs on a lifecycle policy rather than keeping every payload forever.

Budgets are what make any of it operational. Once the spend is split by team or tenant tag, a cost budget per slice with an alert at, say, eighty per cent, plus a forecast-based alert on top, warns each team while they can still act. At a harder threshold a budget action applies a deny IAM policy or an SCP that stops further Bedrock calls, automatically or after someone approves it. Pair each budget with a cost anomaly detection monitor on the same tag dimension, to catch the movements nobody set a threshold for. The untagged remainder in the same view doubles as compliance monitoring for the tagging scheme itself.

Worked example

The bill jumps, and the platform team needs to know who and why before the standup. The estate is carved up by a team tag: on inference profiles for marketing and chat, and on the imported model itself for support. Cost Explorer is the first stop. Filter to the Bedrock service, group by the team tag, and the marketing slice is plainly the one that doubled. One tag key reads the same whichever mechanism put it there, which is what makes a mixed estate reportable.

The “why” needs a finer grain than the tag carries. Marketing’s own profile is shared across several campaigns. The team turns on model invocation logging for that Region and queries the logs in Athena, summing output tokens by the campaign ID their client sets in request metadata on every call. One campaign is generating enormous responses, long output at the dearer output-token rate, which is exactly the shape of a cost spike the token pricing predicts. The fix is a prompt change to cap the response length, plus a switch to batch inference at half the rate for that campaign’s overnight run. Batch jobs run outside the profile, so the team tags the job itself with team=marketing to keep the slice whole.

The loop closes with a Budget. The team sets a cost budget scoped to team=marketing, with an alert at eighty per cent of the monthly allowance and a forecast alert on top. The next campaign that runs hot pages them mid-month instead of surprising finance at the end.

What’s worth remembering

  1. Bedrock on-demand cost is tokens in and tokens out, priced separately with output usually dearer, plus any Provisioned Throughput charged by the model-unit hour whether or not you use it.
  2. Cost allocation tags split the surrounding taggable resources (Lambda, S3, Knowledge Bases) but not a bare InvokeModel call, so tags alone leave the token spend unsplit. Capacity you own is the exception: Provisioned Throughput and imported models have their own ARNs and tags, set at creation, though one resource carries one tag however many tenants call it.
  3. Application inference profiles split shared inference: tag the profile, route a product’s calls through it, and the token cost lands per profile in Cost Explorer and the CUR. Each profile references exactly one foundation model, works only on InvokeModel and Converse, and can never reference an imported model.
  4. IAM principal attribution needs no resource: Bedrock records the calling role on every request, and principal or STS session tags on that identity reach Cost Explorer and CUR 2.0. It is the per-user and per-tenant answer that does not multiply profiles.
  5. Model invocation logging is the only source of per-request token counts, and request metadata puts up to sixteen caller-set key-value pairs in the same record. Both are derived numbers, not billing-grade ones, and the logs cost storage.
  6. AWS Budgets close the loop: a budget per tag with threshold and forecast alerts warns each team before the month ends rather than after, and a budget action can apply a deny IAM policy or an SCP at a harder threshold.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.