Verifying the Exam Room tracks

Every Exam Room track was drafted by agents and published without anyone checking what it claimed. On ANS-C01, checking found about seven wrong facts per post, and every post in both sample batches had at least one. Wrong quotas, inverted rules, a copy-pasteable IAM policy that did not do what the post said. So a track that has not been through this sweep should be assumed wrong, not assumed right.

This file is for picking the work up cold. Nothing here depends on the conversation it started in.

Where things stand

Run it, do not guess:

python3 scripts/exam-verification-status.py            # per-track table
python3 scripts/exam-verification-status.py --todo AIP-C01   # files still to do

State lives in _data/exam_verified.yml, keyed by track: completed for a track that is fully through, verified_slugs for one partway. Do not stamp a verified: field into post frontmatter. An earlier run did, and the extra line pushed nine pop quizzes and flash cards past the 450-word frontmatter ceiling measure-post.py enforces. Publication is not evidence of verification: AIF-C01 and AIP-C01 both finished publishing without ever being checked, which is why the status script counts live and unverified posts separately and flags that combination.

ANS-C01 is done: 113 posts verified, 799 facts corrected, 70 self-contradictions fixed, full coverage against the official exam guide.

AIF-C01, AIP-C01 and CLF-C02 are done too, at zero style warnings and zero rhythm failures. All three have been through the coverage gate, on 13 September 2026:

Track Objectives Covered Gaps Thin Gated Guide version
AIF-C01 69 69 0 0 13 Sep 2026 1.1, 30 Apr 2026
AIP-C01 98 96 0 2 13 Sep 2026 none stated, no revisions page
CLF-C02 135 121 1 13 13 Sep 2026 none stated
AIB-C01 58 56 0 2 17 Sep 2026 none stated, no revisions page
SAA-C03 189 187 0 2 20 Sep 2026 none stated, no revisions page

Weightings were confirmed from the domain pages in every case. Only AIF-C01 publishes a revisions page, at .../ai-practitioner-01/aif-01-revisions.html; AIP-C01, CLF-C02, AIB-C01 and SAA-C03 have no version, date or document history, so there is no cheap way to notice those guides moving and the gate has to be re-run periodically rather than on a revision signal.

AIB-C01 was gated before publication rather than after, which is the order to prefer: a gap means writing posts, and slotting one into a track with nothing live is a scheduling decision rather than a repair. It came back 56/58 with no hard gaps. Both thin objectives are the same shape, and the shape is worth recognising: an objective carried by a four-minute pop quiz plus paragraphs borrowed from posts written for something else, where every neighbouring objective has a 36-43 minute scenario of its own. Skill 2.3.2 (identifying business-model transformation opportunities) teaches the discriminating half well and the generative half not at all; Skill 4.3.5 (workforce development approaches) explains what each instrument produces but never sizes, sequences or measures a programme. Both were closed the same day with one scenario post each, so the track now teaches 58 of 58; the row above records what the gate found, not where the track ended up.

SAA-C03 was planned from the guide before it was written, and the gate shows what that is worth. 187 of 189 objectives taught on the first pass, no hard gaps, two thin, against AIB’s 56 of 58 with two thin on a track a third of the size. Domains 1 and 2 came back 32/32 and 43/43 with nothing thin at all. The difference is the order of work: scripts/workflows/plan-exam-track.js reads the guide domain by domain and hands the writers an objective inventory, so a bullet has to be deliberately dropped rather than quietly missed.

Both thin objectives landed in the same blind spot, which is the one the plan could not see: database choice. Skill 3.3.11 and its Domain 4 twin ask for MySQL compared with PostgreSQL, and every mention of the two in the track is either a migration story or a scattered behavioural fact, so a reader choosing an engine for a new application finds nothing. Skill 4.3.13 asks for cost-effective database types by shape, and the track teaches the columnar half at length and the time-series half not at all, naming Amazon Timestream only to say LiveAnalytics is closed to new customers and never naming Timestream for InfluxDB, which is the option a candidate can still pick. The lesson for the next plan: an objective naming two things (“time series format, columnar format”, “MySQL compared with PostgreSQL”) gets read as one, and the post that answers it teaches whichever half its scenario happened to need.

AIB-C01’s guide stem is ai-business-strategist-01, and the guide carries aib-01-in-scope-services.html, aib-01-out-of-scope-services.html and aib-01-technologies-concepts.html alongside the four domain pages. The gate reads the domain pages only, so the out-of-scope list is worth a separate cheap check: this is the one certification where teaching implementation detail is a defect rather than generosity.

AIB-C01 is done, on 17 September 2026: 57 posts in four batches, 856 documentation lookups, 1,495 claims checked, 346 wrong facts corrected, 6.1 per post, plus 179 arithmetic defects, 73 passages carrying implementation detail the guide puts out of scope, 45 service names that had moved on and 902 style fixes. 8.2M subagent tokens. The whole track finishes at zero style warnings and zero rhythm failures.

Four things that run’s findings are worth carrying to the next track.

Arithmetic is a defect class in its own right, and a read-through misses all of it. 179 across 57 posts, against 346 facts: a scenario states figures, a table repeats them, a worked example computes from them and a takeaway summarises them, and they drift. A population reading 430 in the prose and 508 in the costing table. A revenue ranking stated backwards in the prose, the table and the SVG at once. Give it its own numbered step in the brief and say to follow every number through the post, rather than hoping a careful reader notices.

The dominant staleness is not what the AIP track taught us to expect. On a generative-AI paper it is model IDs and token prices. Here it was serving and lifecycle: a customised Bedrock model was said to have to run on Provisioned Throughput, which stopped being true on 16 July 2025; the twelve-month availability promise was quoted as current when it governs only models launched before 7 September 2026; two posts quoted the AWS CAF Governance Perspective whitepaper, which now carries “This whitepaper is for historical reference only”. A stale framework page is worse than a stale price, because nothing about it looks out of date.

Non-AWS facts were the worst errors in the run, and nothing was checking them. four-thousand-letters-already-sent had the EU AI Act wrong in a way that changes the answer: it called 300 wrong credit decisions a widespread infringement and designed against the two-day Article 73 clock, when Article 3(61) needs harm across at least two Member States and the fitting limb, 3(49)(c), carries fifteen days. It also described UK automated decision-making as the Article 22 restriction, which the Data (Use and Access) Act 2025 replaced on 5 February 2026 with Articles 22A to 22D, and cited two FCA rules at the wrong paragraph. Brief for regulations, standards and jurisdictions explicitly; an agent told to check “AWS facts” will read past all of it. The same post’s scenario had been given a UK priority-services register while set in Australia, where the equivalent is the life-support register under the National Energy Retail Rules.

The currency rule cleared itself. The track carried 193 bare $ figures across 15 posts at the start and zero at the end. Naming the rule in the brief was enough; batch 1 fixed every affected post before anyone had counted them. Worth adding to check-exam-room.py so the next track never accumulates them.

CLF-C02’s findings cluster rather than scatter: five of its six Domain 4 thin items plus its one hard gap sit under Task 4.3 (AWS technical resources and Support options), and three more sit under Task 2.2 in a 30%-weighted domain. Do not re-derive these; the run is in the session transcript and the fixes are tracked separately.

The order of work

Set by which exams Craig is sitting, which beats every other consideration: he is the reader, and he is revising from this material.

  1. CLF-C02. 23 posts, publishing now, and he is studying for it. Done.
  2. AIP-C01. 205 posts, in progress. He is studying for this one too, which is why it gets the full treatment despite being a finished run. The calibration measured 12.1 wrong facts per post here, the worst of any track, because generative-AI content goes stale fastest. Ask the status script for the remaining count rather than trusting a number written here.
  3. Then in publication order, since that is the order he will revise in: ANS-C01 (done), AIB-C01 (January), SAP-C03 (March), DOP-C02 (June), SCS-C03 (September 2027). Each wants verifying before its window opens, while it is still preview: true and has no live URL to protect.
  4. AIF-C01 last, if there is budget. 96 posts, finished publishing on 3 September.

An earlier draft of this file put the live GenAI tracks first on the grounds that wrong content was reaching readers, and ranked a finished run as an archive not worth full verification. That was wrong about who the reader is.

PR and merge in batches of roughly 25 rather than one PR per track. A 200-post PR is unreviewable and one bad file blocks the rest.

Running it

Three workflow scripts live in scripts/workflows/. They are copies: the originals sat under a session directory that dies with its session, and rewriting the prompts from scratch loses the accumulated corrections below.

Exam guides live at https://docs.aws.amazon.com/aws-certification/latest/<stem>/<stem>-domainN.html. Read them; weightings and task statements change between versions.

The stem usually carries a numeric suffix, and guessing it wastes a round of 404s: AIF-C01’s is ai-practitioner-01, not ai-practitioner, and AIB-C01’s is ai-business-strategist-01 rather than anything built from “AI Business”. Probe candidates with curl -o /dev/null -w '%{http_code}' before launching agents, or search for the domain page; do not construct the URL from the certification’s name.

What went wrong before, so it does not again

Never set effort: 'low' on the agent options. A run with it set had all twelve agents blocked from every tool call by a permission-handler error and burned 813k tokens producing nothing. The symptom looks like a schema failure in the summary; the agents’ own transcripts say “every tool call was rejected”.

Judge each file on its own when committing. An early commit guard held the whole batch whenever one file failed a check, so a single flash card stalled everything. Skip files modified in the last 45 seconds (their agent is still writing), skip a file whose frontmatter fails yaml.safe_load or whose check-exam-room.py prints a warning, and commit the rest.

A rhythm failure is usually pre-existing. Flash cards overrun their frontmatter budget on main as well: nine of twelve sampled. Only treat measure-post.py failing as this sweep’s problem if the same file passes on origin/main.

Do not hand-edit a generated checklist. make reading-times rebuilds them from _data/exam_checklists/*.yml. Editing the post directly is overwritten.

On a published post, change content only. Never touch date:, preview:, the filename or the permalink. Re-dating a live post drops it off the listings and the feed while its URL still resolves, and _data/linkedin_shares.yml has public links with UTM parameters attached. scripts/check-published.py catches it; run it before every push.

The branch is deleted when its PR merges. That leaves a stale tracking ref, and git push --force-with-lease then refuses with “stale info” against a branch that no longer exists. git fetch --prune origin fixes it. After a merge, restart from main rather than pushing onto merged history.

A failed restart must not be followed by a push. When the rebase aborted on an unstashed file and the branch was pushed anyway, it stayed based on the old main. The PR went un-mergeable, and GitHub silently declines to run pull_request CI on a conflicted PR, so it presented as a five-hour Actions outage rather than an error. The script now exits non-zero and says so; if it does, resolve with git merge origin/main before pushing. The tell from the GitHub side is mergeable_state: "dirty" on the PR with zero check runs.

Restart the branch with scripts/verify-restart-branch.sh <merged-sha>, not git checkout -B. The agents keep running while CI does, so the commit cycle almost always lands commits after the sha the PR merged, and `checkout -B

origin/main` throws those away without a word. The script rebases them onto the new main instead, and stashes in-flight files across the move because git refuses to rebase over a file an agent is writing. ## SAA-C03 SAA-C03 was verified on 20 September 2026: **all 97 scenario, quiz, card, cheat-sheet and lab posts**, in six batches, every one of them coming back clean. The domain checklist is generated and carries nothing to look up. 1,884 documentation lookups against 4,606 claims produced **557 wrong facts corrected, 6.1 per post**, plus 296 arithmetic defects, 112 passages sitting above or below the associate design level, 40 stale service names and 637 style fixes. The two posts written to close the thin objectives were verified adversarially as they were written, and the checklist is generated. The weekly token limit cut the fifth batch off three posts short. They were finished when it reset, and they were worth the wait: all three are Domain 4, where the arithmetic carries the argument, and they came back at 6.7 facts a post against the track's 6.1. Two of the three had a rejected option rejected for the wrong reason, which reads as sound until somebody checks the number: an AWS WAF rate-based rule ruled out on a floor of 10 that argues nothing at 1,200 a window, when the real objections are the 600-second ceiling on the evaluation window and the documented detection lag; and an API Gateway account throttle called "the limit producing the bill" when at 600 requests a second it is not the binding limit at all. ### What this run says about planning first The coverage gate and the verification pass measure different things, and SAA-C03 is the cleanest demonstration of it in the corpus so far. Planning from the guide took the gate from AIB's 56 of 58 to 187 of 189, and did nothing at all for the error rate: 6.1 facts a post against AIB's 6.1. A track can teach exactly the right objectives and still be wrong about the figures inside them, and no amount of planning catches that. The defect that recurred most in this track was **a post contradicting itself**. One told a reader to raise an X-Ray sampling reservoir on a rule scoped to a child service, then explained in the next sentence why such a rule never fires. Another said a fully pinned RDS Proxy workload keeps connection reuse between sessions, which is the opposite of what pinning does. A third put an EventBridge rule for root sign-in in ap-southeast-2, where that event is never recorded. None of these needs external knowledge to spot; they need somebody to read the post as a whole and follow the claim through. That is what the arithmetic step does for numbers, and it is worth asking the same question of behaviour. ## SAP-C03, complete: the error rate is not the same everywhere 18 of 97 posts, on 20 September 2026, all clean at the end. The number that matters is **257 wrong facts, 14.3 a post**, against 6.1 on SAA-C03 and 6.1 on AIB-C01. Also 76 passages at the wrong level across 18 posts, where SAA-C03 had 112 across 97. So the six-to-eight-a-post figure this file has been quoting is not a property of agent-written posts in general. Two things separate this track. It is the oldest unverified one, drafted as SAP-C02 and retagged, so its claims have had the longest to go stale: 38 stale service names in 18 posts. And it is a professional exam, where a post has more room to be wrong, because the answer turns on mechanism and second-order consequence rather than on which service to pick. Budget a professional track at twice an associate one. ### The finished run confirms it **96 of 96 posts, complete on 28 September 2026. 1,036 wrong facts, 13.3 a post**, plus 283 arithmetic defects, 305 passages at the wrong level, 155 stale service names and 1,055 style fixes, from 1,567 documentation lookups against 2,990 claims. Every post clean at the end. 455 items went into `remaining`. The rate held across every batch rather than spiking in one, which is what makes it the track's property and not a sampling artefact: | Batch | Posts | Facts | Per post | |---|---|---|---| | 1 (2027-05-14 to 05-31) | 18 | 213 | 11.8 | | 2 (2027-06-02 to 06-19) | 18 | 237 | 13.2 | | 3 (2027-06-21 to 07-09) | 18 | 303 | **16.8** | | 4 (2027-07-10 to 07-28) | 18 | 226 | 12.6 | | 5 (labs, 2027-07-30 to 08-04) | 6 | 57 | **9.5** | The spread within the track is worth as much as the average. Batch 3 carried the pricing posts, and pricing is where a drafted post is least reliable: a figure that sounds plausible is wrong in a way no amount of internal consistency catches. Batch 5 was the labs, which came in lowest, because a lab's claims are mostly commands and a command either exists or does not. **The pricing findings from batch 3 are the ones to remember.** One post made two CloudFront claims that were both false. It said CloudFront egress is charged at a lower rate than direct egress from EC2 or S3; for Australian viewers it is identical, USD$0.114/GB, confirmed against the pay-as-you-go page and the Price List API (`APS2-DataTransfer-Out-Bytes`). The rate advantage only exists for offshore viewers hitting US or EU edges at USD$0.085/GB. And it said caching reduces the volume as well as the rate, when CloudFront origin fetches from S3, EC2 and ELB are free, so caching reduces a leg that costs nothing. The same batch had VPC peering offered as cheaper than Transit Gateway on the grounds that peering carries no per-GB charge. It carries USD$0.01/GB each way across zones, the same USD$0.02/GB as a Transit Gateway, and only undercuts a hub when traffic stays inside one zone. S3 RTC's SLA threshold was given as 99.99% in two places, where the SLA commits to 99.9% and 99.99% is the target. Plain CRR was given an invented "99% within 15 minutes under normal load", where AWS's own workload-requirements table says 24 to 48 hours with no SLA at all. **A scenario premise that could not exist**, from batch 2, is the other one worth keeping. The AZ-failure post opened on an ALB with "one subnet mapping"; an ALB must be given at least two subnets in different Availability Zones. The post was describing a configuration no reader could reproduce. Rewritten around two enabled zones with targets in only one, which is the defect that is actually reachable. In the same post RDS Multi-AZ failover was "a minute or so" against AWS's published 60 to 120 seconds, and its worked example completed in 51 — below its own floor. ### The kind of thing a first batch found The RDS cost post's central saving was **impossible**. It moved a 16 TB volume to gp3 and then reduced it to 4 TB; RDS cannot deallocate storage at all ("You can't deallocate space", USER_PIOPS.ModifyingExisting). The same post called the gp2-to-gp3 conversion "the free win", and the Sydney offer file has both at USD$0.138 per GiB-month, so the conversion changes nothing on that invoice. Worse, converting a 4 TiB gp2 volume **halves** baseline throughput, 1,000 MiB/s to 500, so holding throughput constant costs more than it saves. A reader following that post would have done three days of work to raise a bill. A Fargate quiz claimed the EC2 launch type is "the only way to get a GPU under ECS in AWS", which ECS Managed Instances has not been true of since it reached general availability on 30 September 2025. That post's option set predated a launch type AWS now lists among four. Both are the same shape and it is worth naming: **a claim that was true when the post was drafted**. Nothing in the post reads as wrong, the reasoning is sound, and the answer is no longer correct. A read-through cannot catch these, and neither can a coverage gate. Only a lookup does. ## Finishing a batch ``` make reading-times # restamps, regenerates checklists python3 scripts/check-exam-checklists.py --strict python3 scripts/check-exam-room.py <the track's posts> python3 scripts/check-preview-flags.py python3 scripts/check-published.py ``` Then commit, push, open a PR, and merge once CI is green. Add the batch's slugs to `verified_slugs` under its track in `_data/exam_verified.yml`, or the next run will do them again. **Commit as the agents finish, not at the end.** A batch of 25 runs about three at a time, so the slowest agent is an hour behind the first and holding the whole batch uncommitted risks losing all of it. `scripts/verify-commit-cycle.sh ""` commits every settled post and says which it held back and why: a file written in the last 45 seconds, one whose frontmatter will not parse, or one warning in `check-exam-room.py` where the `origin/main` copy does not. Run it on a timer against the workflow journal (`subagents/workflows//journal.jsonl`, one `"type":"result"` line per finished agent) and push once, at the end: every push cancels the in-progress CI run, so a push per cycle means CI never finishes. ## AIB-C01 splits, in progress: what the verifiers found in OTHER posts The sweep of the 29 AIB-C01 split posts is running. Nine are recorded and all nine came back clean on their own text, but seven of them reported a defect in a post they were told not to touch. That is the most valuable output of a sweep over posts that share a scenario, and it is only in the run's journal, so it is written down here before the journal rotates. Each is a claim by a verifier that has not been independently confirmed except where noted. They belong to the parent trim and the parent re-sweep. **Llama 4 Scout's context window on Bedrock is 3.5 million tokens, not 10 million.** `the-handbook-the-assistant-never-read` carries "Llama 4 Scout's model card lists a ten-million-token context window, so there the limit is the meter and not the window", and the figure AWS publishes for Bedrock is 3.5 million. This is the item the writing agents flagged for the sweep: a context-window claim taken off a model card rather than from what Bedrock supports. Its sibling has been corrected and the two now disagree. **Provisioned Throughput Model Unit pricing is published for some families.** `forty-users-then-four-thousand-what-breaks-that-is-not-the-bill` says "AWS publishes neither the throughput a model unit delivers nor its price, directing customers to their account manager". The throughput half is right; the price half is wrong for Nova and Titan, which carry published hourly Model Unit rates in the Price List API. `eleven-years-of-declinatures-and-1-2m-to-spend` has a defensible softer version ("for several model families AWS directs you to your account team for that hourly rate"). Worth knowing because the stronger sentence was quoted in PR #1837 as an example of a post correctly declining to invent a figure, and half of it is an error in the other direction. **A fine-tuning condition on serving a custom model on demand may not exist.** `eleven-years-of-declinatures-and-1-2m-to-spend` claims parameter-efficient versus full-rank fine-tuning decides whether a customised model can be served on demand. AWS's on-demand prerequisites list Region, the 16 July 2025 customisation date, model-access permission and KMS permission, and no such condition. Needs a verdict rather than a silent deletion. **Two Model Monitor dates appear on no AWS page.** `nine-systems-live-and-one-person-watching` states that Model Monitor and Clarify "moved to maintenance on 30 June 2026 and closed to new customers on 30 July 2026". Confirmed directly against `sagemaker/latest/dg/model-monitor.html`: the notice is real and reads "no longer open to new customers. Existing customers can continue to use the service as normal", and it is **undated**. Two invented dates in a track recorded as verified. All 14 posts mentioning Model Monitor do note the closure, so the service's status is handled correctly everywhere; it is only these dates that are wrong. **An adverse-selection claim is inverted, including inside an SVG.** `eleven-years-of-tagging-and-one-customer` reads "the figure is attractive precisely where students take 2.3 titles rather than 4.1" and, in an SVG aria-label, "because the early adopters are the light users". Its own figures refute both (AUD$40.02 a unit, AUD$96 flat, break-even 2.4 titles), so the early adopters are the heavy users. The aria-label matters: a correction that fixes the prose and leaves the accessible text is still wrong for anyone reading it that way. **Two parents resolve a decision their split also resolves.** `three-pilots-to-kill-and-one-to-fund` ends its worked example with three terminate, two pause, four scale, while its split minutes the same review as "continue, reduced scope" and terminates afterwards. Both cannot be true of one review, and every shared figure agrees, so this is the parent still carrying the split's decision. `the-tender-they-lost-on-a-feature-they-dont-have` states an availability baseline as an unqualified deliverable where its split establishes the baseline was never produced. That last pair is the duplication `check-exam-prose.py` found independently by word count, arriving at the same 19 parents from the other direction. A verifier reading for contradictions and a script counting words agreeing on the same posts is the strongest signal available that the parent trim is the right next job. ### The word budget undercounts duplication The 19 parents sent to the trim pass were chosen by `check-exam-prose.py`, on the reasoning that a post carrying two decisions must be over its section budget. That is wrong in one direction, and the verifiers found it. 29 parents spawned a split. Only 19 are over budget. A verifier working on `nine-months-of-prompts-and-nobody-set-a-retention-period` reported that its parent, `ten-percent-cheaper-and-out-of-the-country`, "still carries this post's decision in its own solution paragraph and its takeaway 5", and that parent is **under** the budget and so was never sent for trimming. A post that was already short can carry a second decision compactly and pass a word count. So the budget is a lower bound on duplication, not a detector for it. An earlier note in this file called the budget script and the verifiers "two independent methods landing on the same 19 posts", which overstates it: they agree on those 19, and the verifiers find more. The remaining 10 parents need the same trim, and the only reliable way to find them was to read the pair. What the budget IS good for is ordering the work and proving it finished: it cannot say whether the right half was removed, but it can say a post is still too long, and it cannot be argued with. ### Further cross-post findings from the AIB sweep Each is a verifier's claim about a file it was told not to edit. - `twelve-minutes-saved-and-a-cfo-who-wants-dollars` states "seventy-two case handlers across two sites work about 40,000 cases a quarter, 160,000 a year" while also funding an extension to a second shared-services site, which its split says is not yet staffed. One of the two counts a site twice. - `the-commitment-that-would-have-covered-a-quarter-of-the-bill` says Provisioned Throughput capacity is "sized either in model units or in input and output tokens a minute". AWS's prov-throughput page describes model units only. Check before editing: the Reserved tier does sell tokens a minute, so the sentence may be conflating two products rather than inventing a unit. - `four-systems-four-hundred-thousand-files` describes Amazon Quick's access control as coarse-grained at the knowledge-base level while its split establishes document-level ACLs synced from the source. Flagged as worth a look rather than as a contradiction. - `the-ad-the-model-wrote` lists "prompt attacks" among the harm categories a blurb "clears every one of", which reads as though a prompt-attack filter scores output. It scores input. The same looseness was fixed in its split. - The AWS Marketplace mechanics in `the-contract-under-the-marketplace-listing` (25 accounts, the SCMP addenda, the Vendor Insights counts, variable payments) are repeated in `flash-card-aws-marketplace` and `twelve-tools-and-one- signature`, neither of which has been swept since the split was written. ## AIB-C01 splits, complete: making agents cite their sources did not make them right The 29 split posts were drafted with a step the earlier tracks did not have. The brief made it an instruction rather than advice: before asserting any AWS price, limit, quota, retention maximum, default or behaviour, open the page that states it, read the figure off that page, and log the claim with its URL. The agents complied, returning several hundred cited claims. The sweep says it bought nothing. | | first 57 AIB posts, no cite step | the 29 splits, with it | |---|---|---| | claims checked per post | 26.2 | 31.7 (+21%) | | documentation lookups per post | 15.0 | 15.2 | | **wrong facts per post** | **6.1** | **7.2 (+19%)** | | arithmetic defects per post | 3.1 | 4.6 (+46%) | | above or below the exam's level, per post | 1.3 | 2.7 (+107%) | | service names that had moved on, per post | 0.8 | 0.7 | | **wrong facts per claim checked** | **23.1%** | **22.9%** | The last row is the finding. Accuracy per claim did not move. An agent that opens the AWS page, reads it, and writes the URL down beside what it wrote still gets about one claim in four wrong. What the cite step changed was density: the posts make a fifth more AWS claims, so each one carries a fifth more errors. Two consequences worth holding on to. **A cite step does not substitute for a sweep, and should not be budgeted as if it does.** Expect the same six to eight wrong facts a post whether or not the writer was told to cite. The sweep is where accuracy comes from. **On a non-technical track the cite step has its own failure mode.** Level defects more than doubled, and the mechanism is not mysterious: an agent sent to read AWS documentation reads *implementation* documentation, and then writes implementation detail. On AIB-C01 that is a defect rather than thoroughness, and the brief says so at length. Sending a business-track writer to the service guides pulls it toward the one failure the level note spends most of its words warning about. On a professional track the same step would probably be neutral or useful. What the cite step did produce is a target list. Each writing agent's logged claims told the sweep where to look first, and the two claims flagged at drafting time were both real: **Llama 4 Scout's context window, where AWS publishes two different numbers.** The Bedrock model card lists 10M tokens; the AWS News Blog launch post says "Amazon Bedrock currently supports a 3.5 million token context window for Llama 4 Scout". The post now names both and says which is which, rather than picking one. That is the right handling of a documented disagreement. **Provisioned Throughput comes in two forms, not one.** The purchase page states "Amazon Bedrock offers two types of Provisioned Throughput - by Tokens and by Model Units". Several posts presented Model Units as the only form. The same page also carries a procurement fact the posts had all missed: Model Units must be requested through the AWS support centre before a Provisioned Throughput can be bought, so a reservation is not a same-week decision. On a track where money and timing are the units, that absence mattered more than any rate. ### A fabricated date pair that propagated between services `30 June 2026` and `30 July 2026` appear across the corpus attached to different services, which is the signature of a date that was invented once and then reused. `nine-systems-live-and-one-person-watching` had SageMaker Model Monitor and Clarify "moved to maintenance on 30 June 2026 and closed to new customers on 30 July 2026". The parent trim deleted that passage, for an unrelated reason: its verified split establishes the account never used either service, so neither was selectable. The dates went with it. But the pair survives in `proving-where-ai-content-came-from` (AIP-C01, a track recorded as verified), where `30 July 2026` is now Clarify's closure date and `30 June 2026` is Titan Image Generator G1 v2's end of life. Checked directly, both AWS pages that carry the Model Monitor notice give **no date at all**: `sagemaker/latest/dg/model-monitor.html` and its dedicated `model-monitor-availability-change.html` both read "Amazon SageMaker Model Monitor is no longer open to new customers. Existing customers can continue to use the service as normal", undated. So a closure date for Model Monitor or for Clarify is not something AWS publishes on those pages, and a post giving one is supplying a figure rather than reading it. The Titan Image Generator date may well be real: Bedrock publishes model lifecycle dates on its own page, and that is a different kind of claim from a service closing to new customers. Check it there rather than assuming the pair is wrong together. Worth generalising: when the same date appears against two unrelated services, suspect the date rather than the services. A verifier checking one post cannot see the repetition, so this is a corpus-level check that no per-post sweep will catch. ## The parent trim, complete: 101 contradictions across 30 posts All 30 AIB-C01 parents that spawned a split have been trimmed back to one decision. 30 of 30 returned, none failed, all 30 came back inside the word budget, 74,312 words to 61,050. **101 contradictions resolved.** That is the number worth carrying forward. Each one is a place where a post and its split told a reader different things about the same event, and the splits had been through the documentation sweep while the parents had not, so the split won by default. A sample of what that looked like: - A renewal that one post declines and its split accepts, with different three-year sums behind each. Both cannot be true of one board paper. - A shift-handover summariser already proven at AUD$0.31 a shift in one post and still an open vendor pilot in the other. - One review approving a tariff change "in under two minutes" against the same post's own tier arithmetic, and the split's 485 reviewer-hours a week, both built on three minutes. - A go-live in month six against a February change window owned by another team, which puts first production use in March. - The same AUD$5.9m presented as new revenue in one post and as about minus AUD$500,000 against AUD$6.4m replaced in the other. Most were settled by deleting the parent's version rather than rewriting it, so the track states each fact once, in the post that owns the decision. Two things this says about the method. **A split-and-trim is a verification technique, not just an editing one.** Nobody set out to find 101 contradictions. They surfaced because one post was swept and its twin was not, and an agent was asked to read both and reconcile them. Any pair of posts sharing a scenario can be checked this way, and nothing else in this repository finds this class of defect: a per-post sweep reads one file and cannot see that the file next to it minutes the same meeting differently. **The word budget stops at the parents.** 10 posts are still over, all of them splits, and the growth is not padding: the verifiers added published AWS conditions, and eight of the ten fail on `### Evaluation` alone. The first guess was that the budget counts the side-by-side table as prose, and that is wrong. The table averages 48 words; the worst case carries 390 words of prose against 57 of table. Correcting a table row means explaining the correction beside it, and that prose is the post doing its job. So AIB-C01 is **not** in `_data/prose_budget.yml`, and forcing it in would mean asking an agent to cut prose from ten freshly verified posts, where the thing most likely to go is the AWS condition that made the section long. The budget has done what it can here; the remaining decision is editorial and belongs to whoever next reads those ten. ## AIB-C01 complete: 88 of 88, and what four groups say about accuracy The whole track is swept. 124 post-sweeps across four groups, 3,457 claims checked, 778 wrong facts corrected, 487 arithmetic defects, 259 passages above or below the exam's level. Zero style warnings, zero rhythm failures. Four groups, because the track was written and edited in four distinguishable ways, which makes it the only track that can compare them: | group | cite step | words | facts/post | **per claim** | arith/post | level/post | |---|---|---|---|---|---|---| | first 57, swept 17 Sep | no | ~3,300 | 6.1 | **23.1%** | 3.1 | 1.3 | | 29 splits, written to budget | YES | ~2,100 | 7.2 | **22.9%** | 4.6 | 2.7 | | 19 parents, cut then trimmed | no | ~2,200 | 6.3 | **23.1%** | 4.4 | 2.4 | | 19 parents, cut then trimmed | no | ~2,200 | 5.4 | **19.5%** | 4.8 | 3.4 | **The per-claim column is flat.** Four groups written and edited four different ways, and the rate at which an AWS claim is wrong sits between 19.5% and 23.1% in all of them. Pooled, the 95 posts drafted without a cite step come in at 22.4% and the 29 with it at 22.9%. About one AWS claim in four or five is wrong when a post reaches the sweep, and nothing tried so far moves that. That is the number to plan against. It does not depend on whether the writer was told to open the AWS page, and it does not depend on how heavily the post was edited afterwards. The sweep is where accuracy comes from. **Arithmetic and level defects are a different story, and they went up.** Against the original 57, every group written or edited after the word budget came in carries about 50% more arithmetic defects a post (3.1 to 4.4-4.8) and roughly twice the level defects (1.3 to 2.4-3.4). Two candidate explanations, and the evidence does not separate them cleanly. A deep cut can break arithmetic by removing one operand of a sum that survives elsewhere, which is exactly what the sweep brief was told to look for and what it found. And a shorter post states the same figures in less space, so a tighter-packed post gives arithmetic more chances to disagree with itself per page. The splits were freshly written rather than cut and still show 4.6, which argues for density over cutting damage, but they also make the most claims of any group at 31.7 a post. What it means in practice does not depend on resolving that. **Cutting a track to a word budget raises its arithmetic and level defect density, so sweep after the cut, not before.** Doing it in that order on this track is what surfaced 487 arithmetic defects; sweeping first would have certified prose that was about to change under it, which is the mistake this file already records once.