Verifying a track

One track a week, in a fresh context. This is the whole procedure; the prompt that starts a run is three lines and points here.

EXAM-VERIFICATION.md is the record of what past runs found and what the findings mean. This file is how to do one.

The prompt

Paste this, with the certification code substituted:

Verify the SCS-C03 Exam Room track. Follow VERIFICATION-RUNBOOK.md. Work through as many batches as the session’s budget allows, commit as posts settle, and stop with an honest count rather than a claim of completion.

Nothing else is needed. The level note, the staleness list and the batch of paths all come from the repository.

What one batch looks like

python3 scripts/verify-track-args.py SCS-C03 --status   # how much is left
python3 scripts/verify-track-args.py SCS-C03            # args for the next batch

Feed that JSON to the workflow:

Workflow({ scriptPath: 'scripts/workflows/verify-exam-track-v2.js', args: <the JSON> })

Two workflows at once is the ceiling. The box has four cores and the runner allows CPUs - 2 agents per workflow, so two workflows is four agents and a third only makes all three slower. A post takes most of an hour, so a batch of 18 is about nine hours of wall clock and roughly 3.3M subagent tokens.

While they run, commit what settles:

./scripts/verify-commit-cycle.sh "Verify SCS-C03: more settled posts"
git push

That script holds back a file written in the last 45 seconds, one whose frontmatter will not parse, and one that fails check-exam-room.py where the origin/main copy does not. When it says a file is still being written, wait for it. Committing a half-rewritten post is the one failure this whole arrangement exists to avoid. Push once per cycle rather than per commit: every push cancels the in-progress CI run.

Recording what was done

_data/exam_verified.yml is the only state. After each batch:

python3 scripts/record-verified.py SCS-C03 <the batch's post paths>
python3 scripts/record-verified.py SCS-C03 --complete   # once nothing is left

It takes paths, or --from-journal <run-dir> to read them out of a workflow run’s journal.jsonl. It is idempotent, it refuses a post whose cert_code belongs to another track, and --complete refuses while any post in the track is unrecorded.

verify-track-args.py skips what is recorded, so the next call gives the next batch and a batch that died half way through is picked up by whatever it did manage to record.

Do this even if the session is about to end. A run that verified forty posts and recorded none of them has done the work twice.

Finishing a track

make reading-times                               # restamps, regenerates checklists
python3 scripts/check-exam-checklists.py --strict
python3 scripts/check-preview-flags.py
python3 scripts/check-published.py
python3 scripts/check-currency.py --strict
python3 scripts/check-exam-room.py
python3 scripts/exam-verification-status.py

make reading-times renames nothing but rewrites many files, so run it only once every agent has stopped. The same goes for python3 scripts/relay-exam-queue.py --apply, which renames files and will pull the ground out from under an agent mid-edit.

Then add the track’s row and its findings to EXAM-VERIFICATION.md, open a PR, and merge once CI is green.

What comes back that is not a fix

Two kinds of finding need a person, and the verifier is told to report them rather than act:

A framing built on a dead choice. Where a post’s whole argument is “should we adopt X” and X has closed to new customers, a clause is not the fix. The post either gets retired and its URL redirected to whatever covers the same ground, or rewritten around the service AWS points at instead. Both change a URL, so both are decisions. CloudTrail Lake is the live example: several SCS-C03 posts pick it as the answer, and it closed to new customers on 31 May 2026.

A claim that could not be confirmed. The verifier is told to put these in remaining rather than leave them in the post looking sourced. Read that list at the end of a run; it is usually short and it is usually the most interesting thing in the report.

What to expect

Expect six to eight wrong facts per post on an associate or foundational track, and about twice that on a professional one: SAP-C03 came back at 13.3 across 96 posts against SAA-C03’s 6.1. Within a professional track the pricing posts are the worst (16.8 a post on SAP-C03’s pricing batch) and the labs the mildest (9.5), so a batch’s mix tells you roughly what to expect from it.

Planning fixes coverage and does nothing for accuracy: SAA-C03 gated at 187 of 189 objectives and still carried 6.1 wrong facts a post, the same rate as a track written without a plan. Treat a batch that finds nothing as a weak check rather than a clean batch.

The defect that recurs most is a post contradicting itself a few sentences apart, and it needs no external knowledge to catch: an X-Ray sampling rule scoped to a child service and then explained as a rule that never fires, a claim that a fully pinned RDS Proxy workload keeps connection reuse, an EventBridge rule for root sign-in in a Region where that event is not recorded. The arithmetic step catches the numeric version of this. Ask the same question of behaviour.

Fetching an AWS page

Use scripts/aws-doc.py. A rendered fetch of a docs.aws.amazon.com URL is unreliable in a way that is hard to see: the page renders client-side, so the fetch comes back carrying the service name and nothing else, and the summarising model sometimes fills that gap from its own memory and presents the result as AWS documentation. That is how “Maximum retry attempts: 2, Maximum event age: 3600 seconds” reached an EventBridge claim whose published figures are 185 attempts over 24 hours. A 28-post re-sweep in October 2026 recorded 62 bad fetches, and nearly all of them were this.

The script prefers the markdown twin. Every page under docs.aws.amazon.com is also served at the same path with .md instead of .html, which returns the page as text with no client-side rendering involved, and returns an honest 404 when the page has moved rather than a shell that looks like a page. Where there is no twin, it falls back to curl -L and extracts the text, and it exits non-zero when the page redirected to the guide index, when the body extracts to a shell, or when the text carries {priceOf...} or `` placeholders in place of the figures:

python3 scripts/aws-doc.py <url> --grep '185 times'

A renamed page is resolved from the guide’s own table of contents rather than guessed, which is what turns a 302 into the page you wanted:

python3 scripts/aws-doc.py --find marketplace/latest/buyerguide 'proserv'
  Professional services products -> .../buyer-proserv-products.html

A non-zero exit means go and find the page. It never means soften the claim.

Three cases the script cannot judge, so judge them yourself. A price page’s regional rates are client-rendered, so a figure that arrives with no Region named is the us-east-1 example and an ap-southeast-2 post needs the Price List API instead. Transit Gateway has no offer code of its own, so a NoSuchKey from the Price List API is not “no published price” (its rates sit under AmazonVPC). And a genuinely one-line page exists: Bedrock’s guardrails-supported.html says to go and look at the models table, and that is the whole page.

The re-sweep, and how it survives a cutoff

A sweep’s progress used to live in the running session’s workflow journals, under ~/.claude, which die with the container. A run cut off by a token limit lost its place, and the only record left was a per-track boolean that said “verified” whether the track had been checked to the current standard or a weaker one. The work then either restarted from the top or quietly never finished.

Progress is repo state now. scripts/resweep.py is the only entry point:

python3 scripts/resweep.py --status                     # the backlog, per track
python3 scripts/resweep.py --next SAP-C03 --size 16     # next batch, as workflow args
python3 scripts/resweep.py --record SAP-C03 <paths>     # the moment the batch lands

Feed --next to scripts/workflows/measure-residual.js, then --record that batch before starting the next one. A batch recorded is a batch nobody repeats; a batch left unrecorded is work thrown away when the session ends. --record promotes a track to the current standard when its last post lands, and refuses a path that is not a body post of that track rather than silently dropping it.

Which paths came back

Work out what to record from each agent’s own file field, never from any post path that appears in its output. Both failure directions were measured on one AIF-C01 and SAA-C03 run.

An agent may return an absolute path, /home/user/.../\_posts/2027-03-17-....md, where the rest return repo-relative ones. A pattern anchored on "_posts/ drops it, and the drop is silent: that post sat finished and unrecorded for 85 minutes and would have been re-swept from scratch by whoever picked the track up next.

Reaching instead for any _posts/... path in the agent’s result is worse, and fails in the dangerous direction. An agent’s facts_fixed and remaining prose cites other posts, so a scan of the whole result claimed six posts that no agent had touched, one of them a checklist post that was never in the batch. Recording those would have marked five posts swept that nobody had swept, which is the exact loss the ledger exists to prevent, and nothing downstream would ever have said so.

So: parse the result, read file, resolve it relative to the repo, and report a value that does not resolve rather than skipping it. Then audit, because this class of mistake is invisible from the ledger alone – compare the recorded slugs against the set of file values the run’s agents actually returned, plus whatever was recorded before the run started, and expect nothing left over.

Three things hold the rule up without anyone remembering it. CLAUDE.md carries it, including in the Do Not list. scripts/check-verification-standard.py --strict runs in CI as a ratchet: it never complains about the 694-post backlog, only when the backlog grows, so it cannot be switched off for crying wolf the way a gate on the whole backlog would have been. And a daily Routine, Exam Room re-sweep batch, fires one 16-post batch into a fresh session and stops, which is roughly six weeks to clear the backlog at about 2.6M subagent tokens a firing.

Raising the standard again is the same shape. Add it to standards in _data/exam_verified.yml with what it checks, point current at it, and every track drops to outstanding until re-swept. The ratchet allows that and blocks the reverse.