Exam Room · Advanced Generative AI Developer

Choosing a Model for Code Generation

· 27 min read

Generative AI Development · part of The Exam Room

The situation

A platform team maintains a mid-sized codebase: a few dozen services, a shared internal SDK, and a house style that every new file is expected to follow. They want to add code assistance in several places. Developers want in-editor completions and a chat window grounded in the repo. A migration project needs a service translated from Java to Kotlin. The support rota wants a bot that explains a failing stack trace and drafts a fix. And a reviewer wants a first pass over each pull request that flags the obvious problems before a human looks.

They have two shapes of answer on AWS and they keep conflating them. One is to call a capable foundation model on Amazon Bedrock and build the feature themselves: their own prompts, their own retrieval, their own surface in whatever tool they choose. The other is to adopt the managed product purpose-built for coding, so there is nothing to build. On AWS that product is Kiro, an agentic development environment spanning an IDE, a CLI and a browser interface, built around spec-driven development. It ships a chat pane with the indexed project in context, specs that plan a change into requirements, design and tasks, and agent runs that edit across several files.

The instinct is to pick one for everything. That is the trap. The completions-in-the-editor job and the review-bot-in-the-pipeline job need different things, and one of them is a product you install while the other is a feature you write.

What actually matters

The first thing worth naming is the difference between adopting a product and building a feature. Kiro is finished: it runs in an editor, a terminal and a browser, holds multi-turn chat over the project, plans a change from a written specification, and runs agent tasks across several files. If what you want is developers getting help while they work, that is a subscription and a setup task. Rebuilding it would mean reconstructing a product AWS already ships. If what you want is code assistance inside your own application, a review comment on a pull request, a fix drafted in your support tool, then Kiro’s surfaces are the wrong shape and a model you drive yourself on Bedrock is the fit.

Second is how closely the output has to match your code specifically. A general foundation model handles language mechanics well: idiomatic syntax, common libraries, translating between languages, explaining an error. Your internal SDK, your naming conventions, and the helper every service calls instead of rolling its own are not in its training data. Ungrounded, it emits plausible APIs that do not exist in your repo, which is worse than useless because the result reads as correct. Retrieval fixes that, putting the actual signatures and the actual house patterns in the prompt. Kiro covers the same ground with codebase indexing and steering files, markdown that carries the project’s stack, structure and naming conventions into every request; a custom Bedrock feature needs you to build the retrieval yourself.

Third is context. A single-function completion needs almost none. Reviewing a whole file, translating a service, or reasoning across several modules needs a lot of code in the window at once, so the model’s context length becomes a real constraint. Whole-file and multi-file work requires a model with a large context window; a tight completion does not, and paying for a huge context you never fill is just cost.

Fourth is variability. Code generation usually calls for the expected answer rather than a novel one, so a low temperature is right. Bedrock describes temperature as steepening the distribution over the next token, which makes responses more deterministic. That narrows the spread; it does not pin the output to one string. Explanation and brainstorming tolerate more variety; the code that gets committed should not.

And the axis that outranks the rest, generated code is untrusted until it is checked. A model will produce code that compiles, reads well, and is wrong anyway, or worse, insecure: a hardcoded credential, a SQL string built by concatenation, a dependency with a known vulnerability. Never run generated code unchecked. It goes through the same gates as human-written code, static analysis and dependency scanning, and it is evaluated by running the tests, not by measuring how similar the text looks to a reference. A snippet that scores well on string similarity and fails the test suite is a failure; a snippet that looks nothing like the reference and passes every test is a success.

What we’ll filter on

  1. Build or adopt, are you writing code assistance into your own application, or giving developers an assistant in their editor?
  2. Repo grounding, does the output have to use your real internal APIs and conventions, or is generic-but-correct enough?
  3. Context need, single function, whole file, or across several modules at once?
  4. Variability, does this path need the repeatable expected answer, or room to explore?
  5. Validation, how is the generated code checked, scanned, and test-run before anyone trusts it?

The landscape

A general foundation model on Bedrock. A capable model called through the Bedrock API handles the full spread of code tasks: generation, completion, explanation, translation between languages, and review. You own the prompt, the temperature, the surface, and the integration, which is exactly what you want when the feature lives inside your own application rather than an editor. The cost is that you build everything around the call, and out of the box its output reflects public code rather than your repo.

A foundation model plus retrieval (RAG over the codebase). The same Bedrock model, but the prompt is assembled from a retrieval step that pulls in the relevant real code: the actual function signatures, the internal SDK usage, the house pattern for this kind of file. Now the output names APIs that exist, and new code matches how the codebase already does things. It is more to build, an index over the code and a retrieval step in the request, and it is what makes a custom feature fit your project.

Kiro. The managed-product option: an agentic development environment spanning an IDE, a CLI and a browser, built around spec-driven development, where a specification of requirements, design and tasks is planned first and the agent then works across files. Chat over the indexed project, steering files for house conventions, and multi-file agent runs are what it ships, so a migration runs as a planned job rather than one completion at a time. Check one assumption before adopting it for the editor job: as-you-type ghost-text completion is not in Kiro’s documented feature set. The AWS product that offered that in the IDE was the Amazon Q Developer plugin, closed to new signups on 15 May 2026 and out of support on 30 April 2027, so it is not a foundation for new work. You adopt rather than build, time-to-value is short, and the trade is the usual finished-product one: its surfaces are the ones it ships, a developer’s environment rather than a component inside your own product’s back end.

Large-context models for whole-file and repo work. Within the Bedrock choice, model selection matters for the size of the job. Translating a service or reviewing a full file means holding a lot of code in the window at once, so a model with a large context window is the enabling piece; for single-line completion it is capacity you pay for and never use.

The guardrail and evaluation layer around all of it. Independent of which model or product, generated code passes through review, static analysis, and dependency and secret scanning before it runs, and it is evaluated by executing tests rather than scoring text similarity. Amazon Bedrock Guardrails applies content filters, denied topics, word filters and sensitive-information filters to both the prompt and the response, and its Standard tier extends that detection into code elements: comments, variable and function names, string literals. What it does not do is static analysis, dependency scanning or secret scanning, so those stay with your existing tools. Generated code goes through the checks a human contributor’s code would face.

Evaluation

Side by side

Approach Build or adopt Grounded in your repo Multi-file jobs Own-app surface Ready in-IDE
Bedrock foundation model Build ✗ (generic) Model-dependent ✓ ✗
Bedrock model + codebase RAG Build ✓ Model-dependent ✓ ✗
Kiro Adopt ✓ (indexing, steering) ✓ ✗ ✓
Large-context model on Bedrock Build Via RAG ✓ ✓ ✗

Reading the table against the team’s four jobs: developer-facing help in the editor is the managed product, so adopt Kiro, noting that what it documents is chat and agent runs rather than as-you-type completion. The pull-request review bot and the stack-trace-explaining support bot live inside the team’s own tools, so they are Bedrock features, grounded with retrieval over the repo. The Java-to-Kotlin migration can go either way, a spec-driven agent task in Kiro, or a large-context Bedrock model driven file by file if the migration needs bespoke handling.

The solution

Developer-facing help is the managed product. A chat pane with the indexed project in context, a specification that plans a change into requirements, design and tasks, and agent runs that edit across several files are what Kiro ships. Rebuilding that on raw Bedrock would mean reconstructing an editor integration that already exists as a supported product. Adopt Kiro, point it at the repositories, and standing it up is a setup task rather than a build. Correct one item on the team’s wish list while you are there: ghost-text completion as they type is not documented for Kiro, and the AWS plugin that did it is out of support on 30 April 2027, so treat that as a gap to confirm rather than a given.

The review bot and the support bot are custom Bedrock features. Both live inside the team’s own systems, a comment on a pull request, a reply in the support tool, so there is no editor surface to reuse; the value is in the integration you write. Call a capable model on Bedrock, and ground it with retrieval so the review understands the internal SDK and the fix drafts against APIs that exist. Run these at low temperature, since a code review and a suggested fix call for the expected answer. Put Bedrock Guardrails on the content boundary if the input includes untrusted text; the ApplyGuardrail API evaluates text without invoking a model, so the check can sit either side of the call. Treat the output as untrusted until it has passed the same scanning a human’s code would.

Grounding is what separates useful from plausible. The failure mode of an ungrounded model on a private codebase is plausible fabrication: the output calls client.fetchUser() because that is the common spelling in public code, and your SDK spells it users.get(). Retrieval over the codebase fixes this by putting the real signatures and the real house patterns in front of the model at generation time. Without it, a custom code feature spends its life being corrected; with it, the output fits the project. This is the same lesson as giving the model the right context instead of a longer instruction: the retrieval and the schema around the call decide as much as the wording does.

Validation is test execution, not text similarity. However the code is produced, the gate is the same. Static analysis and dependency and secret scanning catch the insecure patterns, the hardcoded credential, the vulnerable library, that read fine to a human skimming a diff. And quality is measured by running the tests: generated code that passes the suite is good regardless of how little it resembles a reference solution, and code that matches a reference closely but fails a test is not. Any evaluation harness for a code feature runs the code; it does not diff the strings.

Worked example

The team wants a bot that comments on each pull request with a first-pass review before a human looks. It is inside their own pipeline, triggered by the PR event, posting through the code-host API, so it is a Bedrock feature, not an assistant surface.

The naive build sends the diff to a general model with “review this code” and posts whatever comes back. It reads well and it is frequently wrong about this repo: it flags the internal retry helper as a missing error check, because the helper is not in the context, and it suggests a validation call that does not exist in the SDK. Plausible, and useless.

The grounded build assembles the prompt from retrieval. The changed files trigger a lookup that pulls in the real signatures they touch, the internal SDK functions in play, and the house convention for this kind of change, and those go into the context alongside the diff. Now the review is against the code as it actually is: the retry helper is in the context, so the false flag stops, and the suggestions name APIs that exist. The call runs at low temperature, so two runs over the same diff land close together rather than producing two different opinions. If the diff carries untrusted content, a Guardrails policy filters the boundary.

Then the safety gate, which is separate from the model entirely. The bot’s own suggestions, and the human’s code, both pass static analysis and secret and dependency scanning before anyone acts on them; a suggested fix that introduces a concatenated SQL string is caught by the scanner, not trusted because the model produced it. And when the team asks whether the bot is any good, they do not score its comments against a golden review. They take a corpus of pull requests with known issues and measure how many real problems it flags and how many false alarms it raises, running the code where a suggestion is a change. The review bot is evaluated the way the code is: by what happens when you run it.

What’s worth remembering

  1. Adopt Kiro or build on Bedrock. Kiro serves developers in their editor; review bots and support tools inside your own systems are custom Bedrock features.
  2. Kiro documents no as-you-type completion. Chat, specs and multi-file agent runs are documented; ghost-text completion is not, so check the feature set first.
  3. Ground the model in your codebase. Retrieval puts real signatures and house patterns in the prompt, so output uses APIs that exist.
  4. Generated code is untrusted until checked. Apply the same static analysis, dependency and secret scanning as for human code; Bedrock Guardrails does not.
  5. Test by execution, not similarity. Passing the suite is success however little it resembles a reference; matching one but failing a test is not.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.