Verifying a track
One track a week, in a fresh context. This is the whole procedure; the prompt that starts a run is three lines and points here.
EXAM-VERIFICATION.md is the record of what past runs found and what the
findings mean. This file is how to do one.
The prompt
Paste this, with the certification code substituted:
Verify the SCS-C03 Exam Room track. Follow
VERIFICATION-RUNBOOK.md. Work through as many batches as the session’s budget allows, commit as posts settle, and stop with an honest count rather than a claim of completion.
Nothing else is needed. The level note, the staleness list and the batch of paths all come from the repository.
What one batch looks like
python3 scripts/verify-track-args.py SCS-C03 --status # how much is left
python3 scripts/verify-track-args.py SCS-C03 # args for the next batch
Feed that JSON to the workflow:
Workflow({ scriptPath: 'scripts/workflows/verify-exam-track-v2.js', args: <the JSON> })
Two workflows at once is the ceiling. The box has four cores and the runner
allows CPUs - 2 agents per workflow, so two workflows is four agents and a
third only makes all three slower. A post takes most of an hour, so a batch of
18 is about nine hours of wall clock and roughly 3.3M subagent tokens.
While they run, commit what settles:
./scripts/verify-commit-cycle.sh "Verify SCS-C03: more settled posts"
git push
That script holds back a file written in the last 45 seconds, one whose
frontmatter will not parse, and one that fails check-exam-room.py where the
origin/main copy does not. When it says a file is still being written, wait
for it. Committing a half-rewritten post is the one failure this whole
arrangement exists to avoid. Push once per cycle rather than per commit: every
push cancels the in-progress CI run.
Recording what was done
_data/exam_verified.yml is the only state. After each batch:
python3 scripts/record-verified.py SCS-C03 <the batch's post paths>
python3 scripts/record-verified.py SCS-C03 --complete # once nothing is left
It takes paths, or --from-journal <run-dir> to read them out of a workflow
run’s journal.jsonl. It is idempotent, it refuses a post whose cert_code
belongs to another track, and --complete refuses while any post in the track
is unrecorded.
verify-track-args.py skips what is recorded, so the next call gives the next
batch and a batch that died half way through is picked up by whatever it did
manage to record.
Do this even if the session is about to end. A run that verified forty posts and recorded none of them has done the work twice.
Finishing a track
make reading-times # restamps, regenerates checklists
python3 scripts/check-exam-checklists.py --strict
python3 scripts/check-preview-flags.py
python3 scripts/check-published.py
python3 scripts/check-currency.py --strict
python3 scripts/check-exam-room.py
python3 scripts/exam-verification-status.py
make reading-times renames nothing but rewrites many files, so run it only
once every agent has stopped. The same goes for
python3 scripts/relay-exam-queue.py --apply, which renames files and will pull
the ground out from under an agent mid-edit.
Then add the track’s row and its findings to EXAM-VERIFICATION.md, open a PR,
and merge once CI is green.
What comes back that is not a fix
Two kinds of finding need a person, and the verifier is told to report them rather than act:
A framing built on a dead choice. Where a post’s whole argument is “should we adopt X” and X has closed to new customers, a clause is not the fix. The post either gets retired and its URL redirected to whatever covers the same ground, or rewritten around the service AWS points at instead. Both change a URL, so both are decisions. CloudTrail Lake is the live example: several SCS-C03 posts pick it as the answer, and it closed to new customers on 31 May 2026.
A claim that could not be confirmed. The verifier is told to put these in
remaining rather than leave them in the post looking sourced. Read that list
at the end of a run; it is usually short and it is usually the most interesting
thing in the report.
What to expect
Expect six to eight wrong facts per post on an associate or foundational track, and about twice that on a professional one: SAP-C03 came back at 13.3 across 96 posts against SAA-C03’s 6.1. Within a professional track the pricing posts are the worst (16.8 a post on SAP-C03’s pricing batch) and the labs the mildest (9.5), so a batch’s mix tells you roughly what to expect from it.
Planning fixes coverage and does nothing for accuracy: SAA-C03 gated at 187 of 189 objectives and still carried 6.1 wrong facts a post, the same rate as a track written without a plan. Treat a batch that finds nothing as a weak check rather than a clean batch.
The defect that recurs most is a post contradicting itself a few sentences apart, and it needs no external knowledge to catch: an X-Ray sampling rule scoped to a child service and then explained as a rule that never fires, a claim that a fully pinned RDS Proxy workload keeps connection reuse, an EventBridge rule for root sign-in in a Region where that event is not recorded. The arithmetic step catches the numeric version of this. Ask the same question of behaviour.
Fetching an AWS page
Use scripts/aws-doc.py. A rendered fetch of a docs.aws.amazon.com URL is
unreliable in a way that is hard to see: the page renders client-side, so the
fetch comes back carrying the service name and nothing else, and the
summarising model sometimes fills that gap from its own memory and presents the
result as AWS documentation. That is how “Maximum retry attempts: 2, Maximum
event age: 3600 seconds” reached an EventBridge claim whose published figures
are 185 attempts over 24 hours. A 28-post re-sweep in October 2026 recorded 62
bad fetches, and nearly all of them were this.
The script prefers the markdown twin. Every page under docs.aws.amazon.com
is also served at the same path with .md instead of .html, which returns
the page as text with no client-side rendering involved, and returns an honest
404 when the page has moved rather than a shell that looks like a page. Where
there is no twin, it falls back to curl -L and extracts the text, and it
exits non-zero when the page redirected to the guide index, when the body
extracts to a shell, or when the text carries {priceOf...} or ``
placeholders in place of the figures:
python3 scripts/aws-doc.py <url> --grep '185 times'
A renamed page is resolved from the guide’s own table of contents rather than guessed, which is what turns a 302 into the page you wanted:
python3 scripts/aws-doc.py --find marketplace/latest/buyerguide 'proserv'
Professional services products -> .../buyer-proserv-products.html
A non-zero exit means go and find the page. It never means soften the claim.
Three cases the script cannot judge, so judge them yourself. A price page’s
regional rates are client-rendered, so a figure that arrives with no Region
named is the us-east-1 example and an ap-southeast-2 post needs the Price List
API instead. Transit Gateway has no offer code of its own, so a NoSuchKey
from the Price List API is not “no published price” (its rates sit under
AmazonVPC). And a genuinely one-line page exists: Bedrock’s
guardrails-supported.html says to go and look at the models table, and that
is the whole page.
The re-sweep, and how it survives a cutoff
A sweep’s progress used to live in the running session’s workflow journals,
under ~/.claude, which die with the container. A run cut off by a token limit
lost its place, and the only record left was a per-track boolean that said
“verified” whether the track had been checked to the current standard or a
weaker one. The work then either restarted from the top or quietly never
finished.
Progress is repo state now. scripts/resweep.py is the only entry point:
python3 scripts/resweep.py --status # the backlog, per track
python3 scripts/resweep.py --next SAP-C03 --size 16 # next batch, as workflow args
python3 scripts/resweep.py --record SAP-C03 <paths> # the moment the batch lands
Feed --next to scripts/workflows/measure-residual.js, then --record that
batch before starting the next one. A batch recorded is a batch nobody
repeats; a batch left unrecorded is work thrown away when the session ends.
--record promotes a track to the current standard when its last post lands,
and refuses a path that is not a body post of that track rather than silently
dropping it.
Which paths came back
Work out what to record from each agent’s own file field, never from any
post path that appears in its output. Both failure directions were measured on
one AIF-C01 and SAA-C03 run.
An agent may return an absolute path, /home/user/.../\_posts/2027-03-17-....md,
where the rest return repo-relative ones. A pattern anchored on "_posts/ drops
it, and the drop is silent: that post sat finished and unrecorded for 85 minutes
and would have been re-swept from scratch by whoever picked the track up next.
Reaching instead for any _posts/... path in the agent’s result is worse, and
fails in the dangerous direction. An agent’s facts_fixed and remaining prose
cites other posts, so a scan of the whole result claimed six posts that no agent
had touched, one of them a checklist post that was never in the batch. Recording
those would have marked five posts swept that nobody had swept, which is the
exact loss the ledger exists to prevent, and nothing downstream would ever have
said so.
So: parse the result, read file, resolve it relative to the repo, and report a
value that does not resolve rather than skipping it. Then audit, because this
class of mistake is invisible from the ledger alone – compare the recorded
slugs against the set of file values the run’s agents actually returned, plus
whatever was recorded before the run started, and expect nothing left over.
Three things hold the rule up without anyone remembering it. CLAUDE.md
carries it, including in the Do Not list.
scripts/check-verification-standard.py --strict runs in CI as a ratchet: it
never complains about the 694-post backlog, only when the backlog grows, so it
cannot be switched off for crying wolf the way a gate on the whole backlog
would have been. And a daily Routine, Exam Room re-sweep batch, fires one
16-post batch into a fresh session and stops, which is roughly six weeks to
clear the backlog at about 2.6M subagent tokens a firing.
Raising the standard again is the same shape. Add it to standards in
_data/exam_verified.yml with what it checks, point current at it, and every
track drops to outstanding until re-swept. The ratchet allows that and blocks
the reverse.