Exam Room · Advanced Generative AI Developer

Open-Weight or Proprietary: Choosing How You Host a Model

· 37 min read

Generative AI Development · part of The Exam Room

The situation

A product team has two generative-AI features heading for production. The first is a customer-facing assistant that answers billing and account questions in natural language, spiky traffic that peaks during business hours and goes quiet overnight. The second is a batch job that runs every night over a large backlog of long case files, summarising each into a plain-language brief. Its load is steady and predictable. It also needs fine-tuning on the company’s own domain language and house style, so the summaries pass compliance review.

Right now both features call a proprietary foundation model through Amazon Bedrock on demand, billed per input and output token. The assistant is fine that way. The summarisation job is not. The fine-tuning it needs is limited to what the managed model exposes, and the per-token bill on millions of documents a night is climbing fast. Compliance has also started asking whether the model weights and the training data ever leave AWS, and whether the company could keep serving the model if the vendor changed terms.

So the same question lands on both features from opposite directions. One team is content with the managed simplicity it already has. The other needs control a black-box endpoint does not offer, and will run infrastructure to get it. Underneath both is one decision: for this workload, do you call a model someone else operates, or take the weights and host them yourself?

What actually matters

The headline trade is control against managed simplicity. A proprietary model on Bedrock is called as a fully managed, serverless API: no capacity to provision, no scaling to tune, no GPU to keep warm. The catalogue runs past a hundred models, several of them stronger out of the box than anything a team would stand up itself. What you give up is visibility and portability. You cannot see the weights or export them, your customisation is bounded by what the managed service exposes, and the model runs on the vendor’s terms. An open-weight model inverts every one of those. You can download the weights, fine-tune them as deeply as you like, keep them, move them, and inspect what you are running. In return you own the serving infrastructure and everything attached to it.

The second thing that decides the answer is the cost curve, because the two hosting styles bill in different shapes. Managed on-demand inference is per-token: you pay for exactly what you call and nothing while idle, which suits spiky or low-volume traffic. Self-hosting is per-instance-hour, and every hour the endpoint is running is billed whether requests arrive or not. The arithmetic only works once utilisation is high enough that the hourly cost, spread across the tokens served, drops below the per-token rate. A quiet, bursty assistant is cheaper on per-token billing. A saturated, round-the-clock batch job is where a reserved per-hour endpoint gets ahead. Getting this backwards, self-hosting a low-traffic feature or pushing a huge steady load through per-token pricing, is the most common way the bill goes wrong.

Data and weight residency is the third axis, and it is where compliance requirements do the deciding. With a managed proprietary model you control where your prompts and outputs go under the service’s data terms. The weights themselves are never yours and never portable. With an open-weight model you hold the weights and the fine-tuned artefact, so you can keep them inside an account, a Region, or a VPC. You are also not exposed to a vendor changing access or pricing on a model you have built a product around. If the requirement is that the company must be able to keep running this exact model regardless of any vendor, only owning the weights satisfies it.

Then there is licensing, which people skip and later regret. Open-weight does not mean unrestricted. Some open models ship under a permissive licence like Apache 2.0. Others carry community licences with conditions on commercial use above a user threshold, on using outputs to train other models, or on acceptable use. The licence travels with the weights, so read it before a model is built into a product rather than after.

The last two are latency and throughput needs, and the team’s own operational maturity. A self-hosted endpoint lets you pin instance type, Region, and autoscaling policy to a specific latency and throughput target. A shared managed endpoint cannot be tuned that tightly. Reaching for that control assumes a team that can run inference infrastructure: size GPUs, configure autoscaling, patch, monitor, and tune for cost. Handing a model’s operations to a team without the MLOps maturity to carry it is how a self-hosted endpoint becomes a stalled project and a surprise bill. Managed serving exists so a team can skip all of that.

What we’ll filter on

  1. Control and customisation, does the workload need weight access and deep fine-tuning, or is a managed model’s tuning enough?
  2. Cost shape against load, is traffic spiky and low-volume (favours per-token) or steady and high-volume (favours per-hour)?
  3. Data and weight residency, must the weights and fine-tuned artefact stay portable and under your control?
  4. Licensing, does the open model’s licence permit the intended commercial use?
  5. Latency and throughput, does the feature need a pinned, dedicated serving target?
  6. Operational maturity, can the team actually run and optimise inference infrastructure?

The landscape

Proprietary model on Bedrock, on-demand. A fully managed, serverless call to a foundation model whose weights you never see, billed per input and output token. Nothing to provision, scales automatically, and usually the strongest model for the least operational effort. Customisation is bounded by what the service exposes, and you cannot export the model or serve it yourself. This is the default for spiky, low-to-moderate traffic where managed simplicity matters more than control.

Proprietary model on Bedrock, Provisioned Throughput. The same managed model, with capacity reserved and billed hourly per Model unitThe billing block Provisioned Throughput is sold in – one unit delivers a fixed tokens-per-minute rate for a specific model. rather than per token. A unit delivers a set number of input and output tokens a minute, and the hourly rate drops if you commit for one month or six rather than taking the no-commitment option. That gives a fixed throughput level and steadier latency for high, predictable volume, while the model stays a black box. You still cannot see or move the weights; you have changed the billing shape from per-token to per-hour and nothing else. Not every model and Region combination offers Provisioned Throughput, so check the supported list before planning around it.

Open-weight model on Bedrock (managed). Bedrock also serves open-weight models, so you can call one through the same managed, per-token API you use for proprietary models. You get the operational simplicity of Bedrock over an open architecture, with managed fine-tuning where it is offered. You are still calling it as a service rather than holding the weights yourself. Amazon Bedrock Marketplace widens the catalogue to more than a hundred models, but those deploy differently: you subscribe, then deploy to an endpoint hosted by SageMaker AI, choosing the instance type and instance count and paying per instance-hour. Deployment usually takes ten to fifteen minutes, so a Marketplace model is a provisioned endpoint with a Bedrock API in front of it rather than a serverless per-token call.

Bedrock Custom Model Import. Bring weights you have fine-tuned elsewhere into Bedrock’s managed serving path, provided the architecture is one of the supported set. That set covers Llama 2 through 3.3, Mistral, Mixtral, Qwen2 and Qwen3, GPT-OSS, Flan and GPTBigCode, with text weights under 200GB and a maximum context length under 128K. This is the middle ground: you own and customise the weights, and Bedrock runs the serving. Billing is by the Custom Model Units a model copy occupies, charged over five-minute windows from the first successful inference call. If no invocation arrives for five minutes, Bedrock scales the copies to zero and billing stops, and the next call absorbs a cold start measured in tens of seconds depending on model size. Two limits matter before committing: the feature runs in us-east-1, us-east-2, us-west-2 and eu-central-1 only, and an imported model cannot be used with Bedrock’s batch inference API.

Open-weight model self-hosted on SageMaker AI. Deploy an open-weight model to a SageMaker AI real-time endpoint on instances you choose, with autoscaling policies you set, billed per instance-hour. JumpStart carries a catalogue of open models with deploy and fine-tune workflows already wired up, though AWS trimmed that catalogue in March 2026, so confirm a model is still listed. You get full control over fine-tuning, instance type, and serving configuration, and the weights stay in your account. You also own capacity planning, scaling, patching, and cost tuning. Idle instance-hours are billable, with two ways out: asynchronous inference queues requests through S3, handles payloads up to 1GB and runs up to an hour, and scales the instance count to zero between jobs; real-time endpoints scale to zero too, but only when they host inference components, and provisioning back up takes several minutes during which invocations return errors. Serverless inference is not one of the ways out here, because it has no GPU option and caps at 6GB of memory.

Open-weight model self-hosted on EC2. The maximum-control end: run the model on GPU instances you manage directly, with your own serving stack. Total flexibility over every layer, and total responsibility for it, from driver versions to load balancing to keeping the GPUs utilised. Rarely the right first choice unless a requirement genuinely rules out the managed layers above it.

Evaluation

Side by side

Hosting option Weight access and portability Deep fine-tuning Billing shape Ops burden Scaling control Best for
Proprietary on Bedrock, on-demand ✗ ✗ Per token Lowest Automatic Spiky, low-volume, managed simplicity
Proprietary on Bedrock, Provisioned Throughput ✗ ✗ Per model unit, hourly Low Reserved capacity High, steady volume on a black-box model
Open-weight on Bedrock (managed) ✗ Partial (managed) Per token Low Automatic Open architecture, managed serving
Bedrock Custom Model Import ✓ ✓ Per serving capacity Low-medium Managed Your own weights, no endpoint to run
Open-weight on SageMaker AI ✓ ✓ Per instance-hour High You configure Steady load, deep control, weights in-account
Open-weight on EC2 ✓ ✓ Per instance-hour Highest You build it Requirements the managed layers can’t meet

Portability and deep customisation move together in that table, and neither starts until Custom Model Import. Operational burden climbs in step with control. The top three rows give up weight ownership and get managed, mostly per-token serving. The bottom three take on operational effort and get weights you hold on a per-hour cost curve.

Choosing how to host a model on AWS A decision flow from two workload cards, the spiky assistant and the nightly batch job, through three gates asking whether weight access and deep fine-tuning are needed, whether load is steady enough for per-hour billing, and whether the team can run inference infrastructure, to five outcome boxes: Bedrock on-demand, Provisioned Throughput, Custom Model Import, self-hosting on SageMaker AI, and Custom Model Import again as the alternative for a team that skips the endpoint. Spiky assistant bursty, managed-model tuning fine Nightly batch job steady load, needs deep fine-tune Need weight access and deep fine-tuning? Load steady enough for per-hour billing? Team can run inference infrastructure? Proprietary on Bedrock on-demand, per token Provisioned Throughput reserved, per model-unit-hour Custom Model Import your weights, managed serving Self-host, SageMaker AI per instance-hour, in-account Or Custom Model Import no, and traffic spiky no, but load steady yes, own the weights yes no, skip the endpoint

The solution

The spiky assistant should stay a managed proprietary model on Bedrock, on demand. Its traffic is bursty and idle overnight, the shape per-token billing suits: nothing is charged while it is quiet, and it scales through the peak with no capacity to plan. The feature needs neither weight portability nor deep fine-tuning, so the control an open-weight model would add has nothing to do here, and it would leave someone running an endpoint. If the assistant later grows into a high, flat daytime load, the first move is to switch that same model to Provisioned Throughput. That keeps the managed model and changes only the billing shape, steadying the latency and putting a ceiling on the hourly spend.

The nightly summarisation job is the case for leaving the managed on-demand path. It needs fine-tuning deeper than the managed model exposes. Its load is steady and high-volume, so a per-hour endpoint is cheaper than per-token at that utilisation. Compliance requires the weights and the fine-tuned artefact to stay in the account and stay portable regardless of any vendor. That is three of the six filters pointing the same way: control, cost shape, and residency all favour owning the weights. What remains open is how much infrastructure the team is prepared to run.

With the MLOps maturity for it, a self-hosted SageMaker AI endpoint covers everything: right-sized GPU instances, weights fine-tuned on their own case files, autoscaling shaped around the batch window, and asynchronous inference if they want the instance count back at zero between runs. Without it, Custom Model Import is the lighter path, and for a job that only runs at night it can be the cheaper one. Billing accrues in five-minute windows while the batch is running, then stops once five idle minutes pass and the copies scale to zero, so the hours between nightly runs are not billed. A real-time endpoint left standing bills for every one of those hours. The trade is that imported models are not available through Bedrock’s batch inference API, so the job drives them with ordinary InvokeModel calls and its own concurrency control, and the first call after a quiet spell waits out a cold start. Either way the weights stay theirs. The difference is how much of the stack the team holds and what the idle hours cost.

The one filter that overrides all of this is licensing, and it has to be checked before either team commits. An open-weight model is only an option if its licence permits the intended commercial use. A community licence with a user-count threshold or an output-reuse restriction rules out a model that fits on every other axis, and nothing in the deployment flow will flag it. Read the licence attached to the specific weights, because it travels with them into whatever you build.

What changes when the model is an LLM

A self-hosted endpoint here will not behave like the fraud-scoring or classification endpoint a team has run before, and the differences show up in the first week. The shape is still container-based deployment: a model server baked into an image in Amazon ECR, run by a SageMaker AI endpoint, or on ECS or EKS where the team wants the cluster. What differs is container start. The weights load once into GPU memory and stay resident, so a cold start runs to minutes rather than milliseconds, and requests arriving during model loading either queue or time out. Scale-to-zero, the reflex that saves money on a small CPU model, is usually wrong for anything user-facing. Keep a warm floor of at least one instance, stage the weights somewhere fast rather than pulling them over the internet at boot, and treat the endpoint as a long-lived process you are scaling rather than a function you are invoking. The nightly batch job can ignore most of this, because its callers are a scheduler and a queue and both can wait. That is the same reason Custom Model Import’s cold start of tens of seconds does no harm there.

Sizing runs on GPU memory rather than vCPU, and the arithmetic is worth doing on paper before picking an instance. The weights alone are roughly the parameter count multiplied by the bytes per parameter, so an 8-billion-parameter model at 16-bit precision needs about 16GB before anything else is loaded. On top of that sits the KV cache, which grows with both concurrency and sequence length. The cache rather than the weights is what runs an instance out of memory under load: a model that sits comfortably in memory serving one request at a time can fail at thirty concurrent long-context requests on the same hardware. Quantisation to 8-bit or 4-bit gives up a small amount of output quality for a lot of headroom, which lets you either drop to a smaller instance for the same load or run more concurrent sequences on the instance already chosen. For the summarisation job, where inputs are long case files, the cache is the number to size against.

Throughput is measured in token processing capacity rather than requests per second, because two requests to the same model can differ by two orders of magnitude in the work they cause. A server that handles one request at a time leaves the GPU idle for most of every second, which is how a self-hosted endpoint ends up costing more than the per-token bill it replaced. Raising GPU utilisation is the job of continuous batching, where the server admits new sequences into the running batch as older ones finish rather than waiting for the whole batch to complete. That is what the specialised serving frameworks do. On SageMaker AI they arrive as the Large Model Inference containers: one image built around vLLM, another around TensorRT-LLM, both driven by the same configuration format and both supporting continuous batching, quantisation and tensor parallelism. Tune the maximum concurrent sequences and the batch token budget against the GPU memory the weights left spare. Then watch tokens per second, time to first token, and GPU utilisation together, because an endpoint can be saturated on one and idle on another.

Worked example

Put rough numbers on why the batch job flips and the assistant does not. Say the summarisation job processes two million documents a night, several thousand tokens each once the source text and the generated brief are counted. That lands in the billions of tokens a night, every night. On per-token managed pricing that is a large bill that rises linearly with the backlog and never falls away, because the load is constant. A self-hosted endpoint sized for that throughput costs a fixed number of instance-hours a night whether it processes 1.8 or 2.2 million documents. Above the utilisation where the hourly cost spread across the tokens drops below the per-token rate, the endpoint is cheaper, and the gap per document widens as the backlog grows.

The assistant is the mirror image. Its traffic is a few thousand conversations clustered in business hours and almost nothing overnight, so a reserved endpoint would sit idle for most of the day and bill for every hour of it. Per-token pricing charges only for the conversations it actually has, so managed on-demand stays the cheaper shape however the team tunes an endpoint. Same company, same set of hosting styles, opposite answers. The deciding variable is the load curve meeting the billing curve. Fine-tuning depth and weight residency then settle which owned-weight path the batch job takes; utilisation alone settles that it takes one at all.

What’s worth remembering

  1. Control against managed simplicity. Proprietary Bedrock models keep weights hidden and need no operating; open-weight models give you both weights and operations.
  2. Billing shape follows load shape. On-demand bills per token, nothing while idle; self-hosting bills per instance-hour and wins only at high, steady utilisation.
  3. Residency and portability need owned weights. A managed proprietary model offers no way to export the weights or keep serving them yourself.
  4. Custom Model Import: the middle path. Bedrock serves your weights and scales to zero after five idle minutes; imported models cannot use batch inference.
  5. Check the licence first. Open-weight is not licence-free: community licences can limit commercial use above a user threshold, training on outputs, and acceptable use.
  6. Decide per feature. A spiky assistant stays on managed on-demand; a steady, high-volume batch job needing deep fine-tuning moves to owned weights.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.