Verifying the Exam Room tracks
Every Exam Room track was drafted by agents and published without anyone
checking what it claimed. On ANS-C01, checking found about seven wrong facts
per post, and every post in both sample batches had at least one. Wrong
quotas, inverted rules, a copy-pasteable IAM policy that did not do what the
post said. So a track that has not been through this sweep should be assumed
wrong, not assumed right.
This file is for picking the work up cold. Nothing here depends on the
conversation it started in.
Where things stand
Run it, do not guess:
python3 scripts/exam-verification-status.py # per-track table
python3 scripts/exam-verification-status.py --todo AIP-C01 # files still to do
State lives in _data/exam_verified.yml, keyed by track: completed for a
track that is fully through, verified_slugs for one partway. Do not stamp a
verified: field into post frontmatter. An earlier run did, and the extra
line pushed nine pop quizzes and flash cards past the 450-word frontmatter
ceiling measure-post.py enforces. Publication is not evidence of
verification: AIF-C01 and AIP-C01 both finished publishing without ever being
checked, which is why the status script counts live and unverified posts
separately and flags that combination.
ANS-C01 is done: 113 posts verified, 799 facts corrected, 70 self-contradictions
fixed, full coverage against the official exam guide.
AIF-C01, AIP-C01 and CLF-C02 are done too, at zero style warnings and zero
rhythm failures. All three have been through the coverage gate, on
13 September 2026:
| Track |
Objectives |
Covered |
Gaps |
Thin |
Gated |
Guide version |
| AIF-C01 |
69 |
69 |
0 |
0 |
13 Sep 2026 |
1.1, 30 Apr 2026 |
| AIP-C01 |
98 |
96 |
0 |
2 |
13 Sep 2026 |
none stated, no revisions page |
| CLF-C02 |
135 |
121 |
1 |
13 |
13 Sep 2026 |
none stated |
| AIB-C01 |
58 |
56 |
0 |
2 |
17 Sep 2026 |
none stated, no revisions page |
| SAA-C03 |
189 |
187 |
0 |
2 |
20 Sep 2026 |
none stated, no revisions page |
Weightings were confirmed from the domain pages in every case. Only AIF-C01
publishes a revisions page, at .../ai-practitioner-01/aif-01-revisions.html;
AIP-C01, CLF-C02, AIB-C01 and SAA-C03 have no version, date or document history,
so there is no cheap way to notice those guides moving and the gate has to be
re-run periodically rather than on a revision signal.
AIB-C01 was gated before publication rather than after, which is the order to
prefer: a gap means writing posts, and slotting one into a track with nothing
live is a scheduling decision rather than a repair. It came back 56/58 with no
hard gaps. Both thin objectives are the same shape, and the shape is worth
recognising: an objective carried by a four-minute pop quiz plus paragraphs
borrowed from posts written for something else, where every neighbouring
objective has a 36-43 minute scenario of its own. Skill 2.3.2 (identifying
business-model transformation opportunities) teaches the discriminating half
well and the generative half not at all; Skill 4.3.5 (workforce development
approaches) explains what each instrument produces but never sizes, sequences or
measures a programme. Both were closed the same day with one scenario post each, so the track now teaches 58 of 58; the row above records what the gate found, not where the track ended up.
SAA-C03 was planned from the guide before it was written, and the gate shows
what that is worth. 187 of 189 objectives taught on the first pass, no hard
gaps, two thin, against AIB’s 56 of 58 with two thin on a track a third of the
size. Domains 1 and 2 came back 32/32 and 43/43 with nothing thin at all. The
difference is the order of work: scripts/workflows/plan-exam-track.js reads
the guide domain by domain and hands the writers an objective inventory, so a
bullet has to be deliberately dropped rather than quietly missed.
Both thin objectives landed in the same blind spot, which is the one the plan
could not see: database choice. Skill 3.3.11 and its Domain 4 twin ask for
MySQL compared with PostgreSQL, and every mention of the two in the track is
either a migration story or a scattered behavioural fact, so a reader choosing
an engine for a new application finds nothing. Skill 4.3.13 asks for
cost-effective database types by shape, and the track teaches the columnar half
at length and the time-series half not at all, naming Amazon Timestream only to
say LiveAnalytics is closed to new customers and never naming Timestream for
InfluxDB, which is the option a candidate can still pick. The lesson for the
next plan: an objective naming two things (“time series format, columnar
format”, “MySQL compared with PostgreSQL”) gets read as one, and the post that
answers it teaches whichever half its scenario happened to need.
AIB-C01’s guide stem is ai-business-strategist-01, and the guide carries
aib-01-in-scope-services.html, aib-01-out-of-scope-services.html and
aib-01-technologies-concepts.html alongside the four domain pages. The gate
reads the domain pages only, so the out-of-scope list is worth a separate cheap
check: this is the one certification where teaching implementation detail is a
defect rather than generosity.
AIB-C01 is done, on 17 September 2026: 57 posts in four batches, 856
documentation lookups, 1,495 claims checked, 346 wrong facts corrected, 6.1
per post, plus 179 arithmetic defects, 73 passages carrying implementation
detail the guide puts out of scope, 45 service names that had moved on and 902
style fixes. 8.2M subagent tokens. The whole track finishes at zero style
warnings and zero rhythm failures.
Four things that run’s findings are worth carrying to the next track.
Arithmetic is a defect class in its own right, and a read-through misses all
of it. 179 across 57 posts, against 346 facts: a scenario states figures, a
table repeats them, a worked example computes from them and a takeaway
summarises them, and they drift. A population reading 430 in the prose and 508
in the costing table. A revenue ranking stated backwards in the prose, the table
and the SVG at once. Give it its own numbered step in the brief and say to
follow every number through the post, rather than hoping a careful reader
notices.
The dominant staleness is not what the AIP track taught us to expect. On a
generative-AI paper it is model IDs and token prices. Here it was serving and
lifecycle: a customised Bedrock model was said to have to run on Provisioned
Throughput, which stopped being true on 16 July 2025; the twelve-month
availability promise was quoted as current when it governs only models launched
before 7 September 2026; two posts quoted the AWS CAF Governance Perspective
whitepaper, which now carries “This whitepaper is for historical reference
only”. A stale framework page is worse than a stale price, because nothing about
it looks out of date.
Non-AWS facts were the worst errors in the run, and nothing was checking
them. four-thousand-letters-already-sent had the EU AI Act wrong in a way
that changes the answer: it called 300 wrong credit decisions a widespread
infringement and designed against the two-day Article 73 clock, when Article
3(61) needs harm across at least two Member States and the fitting limb,
3(49)(c), carries fifteen days. It also described UK automated decision-making
as the Article 22 restriction, which the Data (Use and Access) Act 2025 replaced
on 5 February 2026 with Articles 22A to 22D, and cited two FCA rules at the
wrong paragraph. Brief for regulations, standards and jurisdictions explicitly;
an agent told to check “AWS facts” will read past all of it. The same post’s
scenario had been given a UK priority-services register while set in Australia,
where the equivalent is the life-support register under the National Energy
Retail Rules.
The currency rule cleared itself. The track carried 193 bare $ figures
across 15 posts at the start and zero at the end. Naming the rule in the brief
was enough; batch 1 fixed every affected post before anyone had counted them.
Worth adding to check-exam-room.py so the next track never accumulates them.
CLF-C02’s findings cluster rather than scatter: five of its six Domain 4 thin
items plus its one hard gap sit under Task 4.3 (AWS technical resources and
Support options), and three more sit under Task 2.2 in a 30%-weighted domain.
Do not re-derive these; the run is in the session transcript and the fixes are
tracked separately.
The order of work
Set by which exams Craig is sitting, which beats every other consideration:
he is the reader, and he is revising from this material.
- CLF-C02. 23 posts, publishing now, and he is studying for it. Done.
- AIP-C01. 205 posts, in progress. He is studying for this one too, which
is why it gets the full treatment despite being a finished run. The
calibration measured 12.1 wrong facts per post here, the worst of any
track, because generative-AI content goes stale fastest. Ask the status
script for the remaining count rather than trusting a number written here.
- Then in publication order, since that is the order he will revise in:
ANS-C01 (done), AIB-C01 (January), SAP-C03 (March), DOP-C02 (June),
SCS-C03 (September 2027). Each wants verifying before its window opens,
while it is still
preview: true and has no live URL to protect.
- AIF-C01 last, if there is budget. 96 posts, finished publishing on
3 September.
An earlier draft of this file put the live GenAI tracks first on the grounds
that wrong content was reaching readers, and ranked a finished run as an
archive not worth full verification. That was wrong about who the reader is.
PR and merge in batches of roughly 25 rather than one PR per track. A 200-post
PR is unreviewable and one bad file blocks the rest.
Running it
Three workflow scripts live in scripts/workflows/. They are copies: the
originals sat under a session directory that dies with its session, and
rewriting the prompts from scratch loses the accumulated corrections below.
verify-exam-track.js is the full treatment, one agent per post: facts
against the docs, then style, rhythm, structure, links and Jekyll. Pass an
array of repo-relative post paths as args. Roughly 161k tokens per post,
measured, which is the number to budget from.
check-internal-consistency.js finds where a post contradicts itself
(a table cell against the prose that argues the case, a worked example
computing from a different number, an SVG description naming a structure other
than the one drawn). No web access, roughly 24k tokens per post. On ANS it
found 70 real defects for an eighth of the price, so it is worth running
across a whole track before committing to full verification.
coverage-gate.js checks every “Knowledge of” and “Skills in” bullet in
the official guide against the track. Read-only. Cheap, four agents, around
600k tokens for a four-domain track. It takes a level that calibrates how
hard a post has to work before it counts as teaching a bullet, and the
calibration is load-bearing: a foundational track judged at professional depth
produces a gap list nobody should act on. business exists for AIB-C01 alone,
where coding and hands-on implementation are out of scope and a post is not
thin for omitting them.
Exam guides live at
https://docs.aws.amazon.com/aws-certification/latest/<stem>/<stem>-domainN.html.
Read them; weightings and task statements change between versions.
The stem usually carries a numeric suffix, and guessing it wastes a round of
404s: AIF-C01’s is ai-practitioner-01, not ai-practitioner, and AIB-C01’s is
ai-business-strategist-01 rather than anything built from “AI Business”.
Probe candidates with curl -o /dev/null -w '%{http_code}' before launching
agents, or search for the domain page; do not construct the URL from the
certification’s name.
What went wrong before, so it does not again
Never set effort: 'low' on the agent options. A run with it set had all
twelve agents blocked from every tool call by a permission-handler error and
burned 813k tokens producing nothing. The symptom looks like a schema failure
in the summary; the agents’ own transcripts say “every tool call was rejected”.
Judge each file on its own when committing. An early commit guard held the
whole batch whenever one file failed a check, so a single flash card stalled
everything. Skip files modified in the last 45 seconds (their agent is still
writing), skip a file whose frontmatter fails yaml.safe_load or whose
check-exam-room.py prints a warning, and commit the rest.
A rhythm failure is usually pre-existing. Flash cards overrun their
frontmatter budget on main as well: nine of twelve sampled. Only treat
measure-post.py failing as this sweep’s problem if the same file passes on
origin/main.
Do not hand-edit a generated checklist. make reading-times rebuilds them
from _data/exam_checklists/*.yml. Editing the post directly is overwritten.
On a published post, change content only. Never touch date:, preview:,
the filename or the permalink. Re-dating a live post drops it off the listings
and the feed while its URL still resolves, and _data/linkedin_shares.yml has
public links with UTM parameters attached. scripts/check-published.py catches
it; run it before every push.
The branch is deleted when its PR merges. That leaves a stale tracking ref,
and git push --force-with-lease then refuses with “stale info” against a
branch that no longer exists. git fetch --prune origin fixes it. After a merge,
restart from main rather than pushing onto merged history.
A failed restart must not be followed by a push. When the rebase aborted on
an unstashed file and the branch was pushed anyway, it stayed based on the old
main. The PR went un-mergeable, and GitHub silently declines to run
pull_request CI on a conflicted PR, so it presented as a five-hour Actions
outage rather than an error. The script now exits non-zero and says so; if it
does, resolve with git merge origin/main before pushing. The tell from the
GitHub side is mergeable_state: "dirty" on the PR with zero check runs.
Restart the branch with scripts/verify-restart-branch.sh <merged-sha>, not
git checkout -B. The agents keep running while CI does, so the commit cycle
almost always lands commits after the sha the PR merged, and `checkout -B
origin/main` throws those away without a word. The script rebases them
onto the new main instead, and stashes in-flight files across the move because
git refuses to rebase over a file an agent is writing.
## SAA-C03
SAA-C03 was verified on 20 September 2026: **all 97 scenario, quiz, card,
cheat-sheet and lab posts**, in six batches, every one of them coming back
clean. The domain checklist is generated and carries nothing to look up. 1,884 documentation lookups against 4,606 claims produced **557 wrong
facts corrected, 6.1 per post**, plus 296 arithmetic defects, 112 passages
sitting above or below the associate design level, 40 stale service names and
637 style fixes. The two posts written to close the thin objectives were
verified adversarially as they were written, and the checklist is generated.
The weekly token limit cut the fifth batch off three posts short. They were
finished when it reset, and they were worth the wait: all three are Domain 4,
where the arithmetic carries the argument, and they came back at 6.7 facts a
post against the track's 6.1. Two of the three had a rejected option rejected
for the wrong reason, which reads as sound until somebody checks the number:
an AWS WAF rate-based rule ruled out on a floor of 10 that argues nothing at
1,200 a window, when the real objections are the 600-second ceiling on the
evaluation window and the documented detection lag; and an API Gateway account
throttle called "the limit producing the bill" when at 600 requests a second
it is not the binding limit at all.
### What this run says about planning first
The coverage gate and the verification pass measure different things, and
SAA-C03 is the cleanest demonstration of it in the corpus so far. Planning
from the guide took the gate from AIB's 56 of 58 to 187 of 189, and did
nothing at all for the error rate: 6.1 facts a post against AIB's 6.1. A
track can teach exactly the right objectives and still be wrong about the
figures inside them, and no amount of planning catches that.
The defect that recurred most in this track was **a post contradicting
itself**. One told a reader to raise an X-Ray sampling reservoir on a rule
scoped to a child service, then explained in the next sentence why such a
rule never fires. Another said a fully pinned RDS Proxy workload keeps
connection reuse between sessions, which is the opposite of what pinning
does. A third put an EventBridge rule for root sign-in in ap-southeast-2,
where that event is never recorded. None of these needs external knowledge to
spot; they need somebody to read the post as a whole and follow the claim
through. That is what the arithmetic step does for numbers, and it is worth
asking the same question of behaviour.
## SAP-C03, complete: the error rate is not the same everywhere
18 of 97 posts, on 20 September 2026, all clean at the end. The number that
matters is **257 wrong facts, 14.3 a post**, against 6.1 on SAA-C03 and 6.1 on
AIB-C01. Also 76 passages at the wrong level across 18 posts, where SAA-C03
had 112 across 97.
So the six-to-eight-a-post figure this file has been quoting is not a property
of agent-written posts in general. Two things separate this track. It is the
oldest unverified one, drafted as SAP-C02 and retagged, so its claims have had
the longest to go stale: 38 stale service names in 18 posts. And it is a
professional exam, where a post has more room to be wrong, because the answer
turns on mechanism and second-order consequence rather than on which service
to pick.
Budget a professional track at twice an associate one.
### The finished run confirms it
**96 of 96 posts, complete on 28 September 2026. 1,036 wrong facts, 13.3 a
post**, plus 283 arithmetic defects, 305 passages at the wrong level, 155 stale
service names and 1,055 style fixes, from 1,567 documentation lookups against
2,990 claims. Every post clean at the end. 455 items went into `remaining`.
The rate held across every batch rather than spiking in one, which is what makes
it the track's property and not a sampling artefact:
| Batch | Posts | Facts | Per post |
|---|---|---|---|
| 1 (2027-05-14 to 05-31) | 18 | 213 | 11.8 |
| 2 (2027-06-02 to 06-19) | 18 | 237 | 13.2 |
| 3 (2027-06-21 to 07-09) | 18 | 303 | **16.8** |
| 4 (2027-07-10 to 07-28) | 18 | 226 | 12.6 |
| 5 (labs, 2027-07-30 to 08-04) | 6 | 57 | **9.5** |
The spread within the track is worth as much as the average. Batch 3 carried the
pricing posts, and pricing is where a drafted post is least reliable: a figure
that sounds plausible is wrong in a way no amount of internal consistency
catches. Batch 5 was the labs, which came in lowest, because a lab's claims are
mostly commands and a command either exists or does not.
**The pricing findings from batch 3 are the ones to remember.** One post made
two CloudFront claims that were both false. It said CloudFront egress is charged
at a lower rate than direct egress from EC2 or S3; for Australian viewers it is
identical, USD$0.114/GB, confirmed against the pay-as-you-go page and the Price
List API (`APS2-DataTransfer-Out-Bytes`). The rate advantage only exists for
offshore viewers hitting US or EU edges at USD$0.085/GB. And it said caching
reduces the volume as well as the rate, when CloudFront origin fetches from S3,
EC2 and ELB are free, so caching reduces a leg that costs nothing.
The same batch had VPC peering offered as cheaper than Transit Gateway on the
grounds that peering carries no per-GB charge. It carries USD$0.01/GB each way
across zones, the same USD$0.02/GB as a Transit Gateway, and only undercuts a
hub when traffic stays inside one zone. S3 RTC's SLA threshold was given as
99.99% in two places, where the SLA commits to 99.9% and 99.99% is the target.
Plain CRR was given an invented "99% within 15 minutes under normal load", where
AWS's own workload-requirements table says 24 to 48 hours with no SLA at all.
**A scenario premise that could not exist**, from batch 2, is the other one
worth keeping. The AZ-failure post opened on an ALB with "one subnet mapping";
an ALB must be given at least two subnets in different Availability Zones. The
post was describing a configuration no reader could reproduce. Rewritten around
two enabled zones with targets in only one, which is the defect that is actually
reachable. In the same post RDS Multi-AZ failover was "a minute or so" against
AWS's published 60 to 120 seconds, and its worked example completed in 51 —
below its own floor.
### The kind of thing a first batch found
The RDS cost post's central saving was **impossible**. It moved a 16 TB volume
to gp3 and then reduced it to 4 TB; RDS cannot deallocate storage at all
("You can't deallocate space", USER_PIOPS.ModifyingExisting). The same post
called the gp2-to-gp3 conversion "the free win", and the Sydney offer file has
both at USD$0.138 per GiB-month, so the conversion changes nothing on that
invoice. Worse, converting a 4 TiB gp2 volume **halves** baseline throughput,
1,000 MiB/s to 500, so holding throughput constant costs more than it saves. A
reader following that post would have done three days of work to raise a bill.
A Fargate quiz claimed the EC2 launch type is "the only way to get a GPU under
ECS in AWS", which ECS Managed Instances has not been true of since it reached
general availability on 30 September 2025. That post's option set predated a
launch type AWS now lists among four.
Both are the same shape and it is worth naming: **a claim that was true when
the post was drafted**. Nothing in the post reads as wrong, the reasoning is
sound, and the answer is no longer correct. A read-through cannot catch these,
and neither can a coverage gate. Only a lookup does.
## Finishing a batch
```
make reading-times # restamps, regenerates checklists
python3 scripts/check-exam-checklists.py --strict
python3 scripts/check-exam-room.py <the track's posts>
python3 scripts/check-preview-flags.py
python3 scripts/check-published.py
```
Then commit, push, open a PR, and merge once CI is green. Add the batch's slugs
to `verified_slugs` under its track in `_data/exam_verified.yml`, or the next
run will do them again.
**Commit as the agents finish, not at the end.** A batch of 25 runs about three
at a time, so the slowest agent is an hour behind the first and holding the
whole batch uncommitted risks losing all of it.
`scripts/verify-commit-cycle.sh ""` commits every settled post and
says which it held back and why: a file written in the last 45 seconds, one
whose frontmatter will not parse, or one warning in `check-exam-room.py` where
the `origin/main` copy does not. Run it on a timer against the workflow journal
(`subagents/workflows//journal.jsonl`, one `"type":"result"` line per
finished agent) and push once, at the end: every push cancels the in-progress
CI run, so a push per cycle means CI never finishes.
## AIB-C01 splits, in progress: what the verifiers found in OTHER posts
The sweep of the 29 AIB-C01 split posts is running. Nine are recorded and all
nine came back clean on their own text, but seven of them reported a defect in
a post they were told not to touch. That is the most valuable output of a sweep
over posts that share a scenario, and it is only in the run's journal, so it is
written down here before the journal rotates.
Each is a claim by a verifier that has not been independently confirmed except
where noted. They belong to the parent trim and the parent re-sweep.
**Llama 4 Scout's context window on Bedrock is 3.5 million tokens, not 10
million.** `the-handbook-the-assistant-never-read` carries "Llama 4 Scout's
model card lists a ten-million-token context window, so there the limit is the
meter and not the window", and the figure AWS publishes for Bedrock is 3.5
million. This is the item the writing agents flagged for the sweep: a
context-window claim taken off a model card rather than from what Bedrock
supports. Its sibling has been corrected and the two now disagree.
**Provisioned Throughput Model Unit pricing is published for some families.**
`forty-users-then-four-thousand-what-breaks-that-is-not-the-bill` says "AWS
publishes neither the throughput a model unit delivers nor its price, directing
customers to their account manager". The throughput half is right; the price
half is wrong for Nova and Titan, which carry published hourly Model Unit rates
in the Price List API. `eleven-years-of-declinatures-and-1-2m-to-spend` has a
defensible softer version ("for several model families AWS directs you to your
account team for that hourly rate"). Worth knowing because the stronger
sentence was quoted in PR #1837 as an example of a post correctly declining to
invent a figure, and half of it is an error in the other direction.
**A fine-tuning condition on serving a custom model on demand may not exist.**
`eleven-years-of-declinatures-and-1-2m-to-spend` claims parameter-efficient
versus full-rank fine-tuning decides whether a customised model can be served
on demand. AWS's on-demand prerequisites list Region, the 16 July 2025
customisation date, model-access permission and KMS permission, and no such
condition. Needs a verdict rather than a silent deletion.
**Two Model Monitor dates appear on no AWS page.**
`nine-systems-live-and-one-person-watching` states that Model Monitor and
Clarify "moved to maintenance on 30 June 2026 and closed to new customers on
30 July 2026". Confirmed directly against
`sagemaker/latest/dg/model-monitor.html`: the notice is real and reads "no
longer open to new customers. Existing customers can continue to use the
service as normal", and it is **undated**. Two invented dates in a track
recorded as verified. All 14 posts mentioning Model Monitor do note the
closure, so the service's status is handled correctly everywhere; it is only
these dates that are wrong.
**An adverse-selection claim is inverted, including inside an SVG.**
`eleven-years-of-tagging-and-one-customer` reads "the figure is attractive
precisely where students take 2.3 titles rather than 4.1" and, in an SVG
aria-label, "because the early adopters are the light users". Its own figures
refute both (AUD$40.02 a unit, AUD$96 flat, break-even 2.4 titles), so the
early adopters are the heavy users. The aria-label matters: a correction that
fixes the prose and leaves the accessible text is still wrong for anyone
reading it that way.
**Two parents resolve a decision their split also resolves.**
`three-pilots-to-kill-and-one-to-fund` ends its worked example with three
terminate, two pause, four scale, while its split minutes the same review as
"continue, reduced scope" and terminates afterwards. Both cannot be true of one
review, and every shared figure agrees, so this is the parent still carrying
the split's decision. `the-tender-they-lost-on-a-feature-they-dont-have` states
an availability baseline as an unqualified deliverable where its split
establishes the baseline was never produced.
That last pair is the duplication `check-exam-prose.py` found independently by
word count, arriving at the same 19 parents from the other direction. A
verifier reading for contradictions and a script counting words agreeing on the
same posts is the strongest signal available that the parent trim is the right
next job.
### The word budget undercounts duplication
The 19 parents sent to the trim pass were chosen by `check-exam-prose.py`, on the
reasoning that a post carrying two decisions must be over its section budget.
That is wrong in one direction, and the verifiers found it.
29 parents spawned a split. Only 19 are over budget. A verifier working on
`nine-months-of-prompts-and-nobody-set-a-retention-period` reported that its
parent, `ten-percent-cheaper-and-out-of-the-country`, "still carries this post's
decision in its own solution paragraph and its takeaway 5", and that parent is
**under** the budget and so was never sent for trimming. A post that was already
short can carry a second decision compactly and pass a word count.
So the budget is a lower bound on duplication, not a detector for it. An earlier
note in this file called the budget script and the verifiers "two independent
methods landing on the same 19 posts", which overstates it: they agree on those
19, and the verifiers find more. The remaining 10 parents need the same trim,
and the only reliable way to find them was to read the pair.
What the budget IS good for is ordering the work and proving it finished: it
cannot say whether the right half was removed, but it can say a post is still
too long, and it cannot be argued with.
### Further cross-post findings from the AIB sweep
Each is a verifier's claim about a file it was told not to edit.
- `twelve-minutes-saved-and-a-cfo-who-wants-dollars` states "seventy-two case
handlers across two sites work about 40,000 cases a quarter, 160,000 a year"
while also funding an extension to a second shared-services site, which its
split says is not yet staffed. One of the two counts a site twice.
- `the-commitment-that-would-have-covered-a-quarter-of-the-bill` says
Provisioned Throughput capacity is "sized either in model units or in input
and output tokens a minute". AWS's prov-throughput page describes model units
only. Check before editing: the Reserved tier does sell tokens a minute, so
the sentence may be conflating two products rather than inventing a unit.
- `four-systems-four-hundred-thousand-files` describes Amazon Quick's access
control as coarse-grained at the knowledge-base level while its split
establishes document-level ACLs synced from the source. Flagged as worth a
look rather than as a contradiction.
- `the-ad-the-model-wrote` lists "prompt attacks" among the harm categories a
blurb "clears every one of", which reads as though a prompt-attack filter
scores output. It scores input. The same looseness was fixed in its split.
- The AWS Marketplace mechanics in `the-contract-under-the-marketplace-listing`
(25 accounts, the SCMP addenda, the Vendor Insights counts, variable payments)
are repeated in `flash-card-aws-marketplace` and `twelve-tools-and-one-
signature`, neither of which has been swept since the split was written.
## AIB-C01 splits, complete: making agents cite their sources did not make them right
The 29 split posts were drafted with a step the earlier tracks did not have. The
brief made it an instruction rather than advice: before asserting any AWS price,
limit, quota, retention maximum, default or behaviour, open the page that states
it, read the figure off that page, and log the claim with its URL. The agents
complied, returning several hundred cited claims.
The sweep says it bought nothing.
| | first 57 AIB posts, no cite step | the 29 splits, with it |
|---|---|---|
| claims checked per post | 26.2 | 31.7 (+21%) |
| documentation lookups per post | 15.0 | 15.2 |
| **wrong facts per post** | **6.1** | **7.2 (+19%)** |
| arithmetic defects per post | 3.1 | 4.6 (+46%) |
| above or below the exam's level, per post | 1.3 | 2.7 (+107%) |
| service names that had moved on, per post | 0.8 | 0.7 |
| **wrong facts per claim checked** | **23.1%** | **22.9%** |
The last row is the finding. Accuracy per claim did not move. An agent that
opens the AWS page, reads it, and writes the URL down beside what it wrote still
gets about one claim in four wrong. What the cite step changed was density: the
posts make a fifth more AWS claims, so each one carries a fifth more errors.
Two consequences worth holding on to.
**A cite step does not substitute for a sweep, and should not be budgeted as if
it does.** Expect the same six to eight wrong facts a post whether or not the
writer was told to cite. The sweep is where accuracy comes from.
**On a non-technical track the cite step has its own failure mode.** Level
defects more than doubled, and the mechanism is not mysterious: an agent sent to
read AWS documentation reads *implementation* documentation, and then writes
implementation detail. On AIB-C01 that is a defect rather than thoroughness, and
the brief says so at length. Sending a business-track writer to the service
guides pulls it toward the one failure the level note spends most of its words
warning about. On a professional track the same step would probably be neutral
or useful.
What the cite step did produce is a target list. Each writing agent's logged
claims told the sweep where to look first, and the two claims flagged at
drafting time were both real:
**Llama 4 Scout's context window, where AWS publishes two different numbers.**
The Bedrock model card lists 10M tokens; the AWS News Blog launch post says
"Amazon Bedrock currently supports a 3.5 million token context window for Llama
4 Scout". The post now names both and says which is which, rather than picking
one. That is the right handling of a documented disagreement.
**Provisioned Throughput comes in two forms, not one.** The purchase page states
"Amazon Bedrock offers two types of Provisioned Throughput - by Tokens and by
Model Units". Several posts presented Model Units as the only form. The same page
also carries a procurement fact the posts had all missed: Model Units must be
requested through the AWS support centre before a Provisioned Throughput can be
bought, so a reservation is not a same-week decision. On a track where money and
timing are the units, that absence mattered more than any rate.
### A fabricated date pair that propagated between services
`30 June 2026` and `30 July 2026` appear across the corpus attached to different
services, which is the signature of a date that was invented once and then
reused.
`nine-systems-live-and-one-person-watching` had SageMaker Model Monitor and
Clarify "moved to maintenance on 30 June 2026 and closed to new customers on
30 July 2026". The parent trim deleted that passage, for an unrelated reason:
its verified split establishes the account never used either service, so neither
was selectable. The dates went with it.
But the pair survives in `proving-where-ai-content-came-from` (AIP-C01, a track
recorded as verified), where `30 July 2026` is now Clarify's closure date and
`30 June 2026` is Titan Image Generator G1 v2's end of life.
Checked directly, both AWS pages that carry the Model Monitor notice give **no
date at all**: `sagemaker/latest/dg/model-monitor.html` and its dedicated
`model-monitor-availability-change.html` both read "Amazon SageMaker Model
Monitor is no longer open to new customers. Existing customers can continue to
use the service as normal", undated. So a closure date for Model Monitor or for
Clarify is not something AWS publishes on those pages, and a post giving one is
supplying a figure rather than reading it.
The Titan Image Generator date may well be real: Bedrock publishes model
lifecycle dates on its own page, and that is a different kind of claim from a
service closing to new customers. Check it there rather than assuming the pair
is wrong together.
Worth generalising: when the same date appears against two unrelated services,
suspect the date rather than the services. A verifier checking one post cannot
see the repetition, so this is a corpus-level check that no per-post sweep will
catch.
## The parent trim, complete: 101 contradictions across 30 posts
All 30 AIB-C01 parents that spawned a split have been trimmed back to one
decision. 30 of 30 returned, none failed, all 30 came back inside the word
budget, 74,312 words to 61,050.
**101 contradictions resolved.** That is the number worth carrying forward. Each
one is a place where a post and its split told a reader different things about
the same event, and the splits had been through the documentation sweep while
the parents had not, so the split won by default. A sample of what that looked
like:
- A renewal that one post declines and its split accepts, with different
three-year sums behind each. Both cannot be true of one board paper.
- A shift-handover summariser already proven at AUD$0.31 a shift in one post and
still an open vendor pilot in the other.
- One review approving a tariff change "in under two minutes" against the same
post's own tier arithmetic, and the split's 485 reviewer-hours a week, both
built on three minutes.
- A go-live in month six against a February change window owned by another team,
which puts first production use in March.
- The same AUD$5.9m presented as new revenue in one post and as about minus
AUD$500,000 against AUD$6.4m replaced in the other.
Most were settled by deleting the parent's version rather than rewriting it, so
the track states each fact once, in the post that owns the decision.
Two things this says about the method.
**A split-and-trim is a verification technique, not just an editing one.** Nobody
set out to find 101 contradictions. They surfaced because one post was swept and
its twin was not, and an agent was asked to read both and reconcile them. Any
pair of posts sharing a scenario can be checked this way, and nothing else in
this repository finds this class of defect: a per-post sweep reads one file and
cannot see that the file next to it minutes the same meeting differently.
**The word budget stops at the parents.** 10 posts are still over, all of them
splits, and the growth is not padding: the verifiers added published AWS
conditions, and eight of the ten fail on `### Evaluation` alone. The first guess
was that the budget counts the side-by-side table as prose, and that is wrong.
The table averages 48 words; the worst case carries 390 words of prose against
57 of table. Correcting a table row means explaining the correction beside it,
and that prose is the post doing its job.
So AIB-C01 is **not** in `_data/prose_budget.yml`, and forcing it in would mean
asking an agent to cut prose from ten freshly verified posts, where the thing
most likely to go is the AWS condition that made the section long. The budget
has done what it can here; the remaining decision is editorial and belongs to
whoever next reads those ten.
## AIB-C01 complete: 88 of 88, and what four groups say about accuracy
The whole track is swept. 124 post-sweeps across four groups, 3,457 claims
checked, 778 wrong facts corrected, 487 arithmetic defects, 259 passages above or
below the exam's level. Zero style warnings, zero rhythm failures.
Four groups, because the track was written and edited in four distinguishable
ways, which makes it the only track that can compare them:
| group | cite step | words | facts/post | **per claim** | arith/post | level/post |
|---|---|---|---|---|---|---|
| first 57, swept 17 Sep | no | ~3,300 | 6.1 | **23.1%** | 3.1 | 1.3 |
| 29 splits, written to budget | YES | ~2,100 | 7.2 | **22.9%** | 4.6 | 2.7 |
| 19 parents, cut then trimmed | no | ~2,200 | 6.3 | **23.1%** | 4.4 | 2.4 |
| 19 parents, cut then trimmed | no | ~2,200 | 5.4 | **19.5%** | 4.8 | 3.4 |
**The per-claim column is flat.** Four groups written and edited four different
ways, and the rate at which an AWS claim is wrong sits between 19.5% and 23.1%
in all of them. Pooled, the 95 posts drafted without a cite step come in at
22.4% and the 29 with it at 22.9%. About one AWS claim in four or five is wrong
when a post reaches the sweep, and nothing tried so far moves that.
That is the number to plan against. It does not depend on whether the writer was
told to open the AWS page, and it does not depend on how heavily the post was
edited afterwards. The sweep is where accuracy comes from.
**Arithmetic and level defects are a different story, and they went up.** Against
the original 57, every group written or edited after the word budget came in
carries about 50% more arithmetic defects a post (3.1 to 4.4-4.8) and roughly
twice the level defects (1.3 to 2.4-3.4).
Two candidate explanations, and the evidence does not separate them cleanly. A
deep cut can break arithmetic by removing one operand of a sum that survives
elsewhere, which is exactly what the sweep brief was told to look for and what it
found. And a shorter post states the same figures in less space, so a
tighter-packed post gives arithmetic more chances to disagree with itself per
page. The splits were freshly written rather than cut and still show 4.6, which
argues for density over cutting damage, but they also make the most claims of any
group at 31.7 a post.
What it means in practice does not depend on resolving that. **Cutting a track to
a word budget raises its arithmetic and level defect density, so sweep after the
cut, not before.** Doing it in that order on this track is what surfaced 487
arithmetic defects; sweeping first would have certified prose that was about to
change under it, which is the mistake this file already records once.