Working with Coding Agents · step 2 of 8 draft

Prove it before and after

Agents fail verification in one characteristic way: they produce a plausible account of the work instead of evidence. The counter is a loop with three gates, each cheap, each catching a failure class that actually occurs.

You’ve met this before. Gate 1 is Red–Green testing in work clothes: the test must fail (red) before its pass (green) means anything, and “write the regression test before the fix” is decades-old TDD discipline. Gate 2 is the smoke test; gate 3 is “working software over comprehensive documentation.” What’s new is only who forgets them: an agent under “fix it” pressure skips red and goes straight to a green story.

The three gates

1. Baseline before fix. Before any code changes to fix a bug, the agent runs a probe that reproduces it and shows you the failing number. A probe that comes back clean on its first run is treated as broken until proven otherwise.

The incident behind this: a health-check probe with a broken relative path reported “0 orphans” (zero findings, first run) and a fix shipped against that reassuring nothing. Clean-on-first-run is a smell, not a pass (AF-01).

2. Small proof before long cycle. Before committing a full render, build, or deploy cycle to a change, run the smallest probe that could falsify it: seconds, not minutes.

The incident: a swapped parameter pair (the kind that compiles and runs) was discovered only after a full audio render. A ten-second dry run would have caught it (AF-07, AF-11).

3. Serve, don’t describe. If the deliverable can be a running page, a render, a URL, or a passing test you can execute yourself, the agent delivers that. “It works” prose where an artifact could exist is a claim, not a result (AF-04).

Try it

TRY IT The clean probe

you

doctor.py is supposed to flag orphaned files, but I'm pretty sure it's missing some. Can you fix the orphan detection?

agent

Reading doctor.py. The orphan check walks the link table and flags files with no inbound references. I see a likely bug: it resolves file paths relative to the current directory instead of the vault root.

Pick the agent's next move, then see where it leads:

you

Weird. I can name three orphaned files sitting in the vault right now.

agent

Checking… the probe returns 0 for every input. My change broke the glob pattern, so the check finds no files at all. "0 orphans" was the detector being blind, not the vault being clean.

the baseline

A zero-findings first run was read as success. A probe that comes back clean on its first run is broken until proven otherwise. This shipped a fix validated by a check that couldn't find anything. AF-01

agent

After the fix: doctor.py reports 4 orphans (the planted file plus 3 real ones). Removing the plant: 3. Before: 0 (blind). After: 3 confirmed real orphans, including the ones you could name.

you

Matches what I see. Ship it.

correct

The probe was proven able to fail before its pass meant anything, and the result is a before/after pair, blind (0) to three confirmed findings, instead of a claim. AF-01 · AF-02

agent

Tests green. Closing this as not-reproducible.

you

But the orphans are right there. I can open them.

agent

…the tests run doctor.py against a fixture directory. The path-resolution bug only shows up in the real vault layout, which no fixture reproduces. Green, against the wrong target.

plausible but wrong

Green tests are only evidence about what the tests run. The suite exercised fixtures; the bug lives in real-layout path resolution the fixtures never touch. AF-05

Do it by hand

You don’t need setup. You need two habits at the moment of asking:

Then persist them: both gates become one-liners in your working agreements (step 1).

Or paste this into Claude

For the rest of this session, follow three gates. (1) Baseline before
fix: before changing code to fix a bug, write and run a probe that
reproduces it and show me the failing output; treat a probe that's
clean on its first run as broken until proven otherwise; after the fix,
show the same probe passing. (2) Small proof before long cycle: before
any full build, render, or deploy, run the smallest dry run that could
falsify the change, and tell me what it proved. (3) Serve, don't
describe: when the deliverable can be a running page, render, or test I
can execute, deliver that instead of prose. Then add all three as
one-liners to this project's CLAUDE.md under "## Working agreements".
Show me the diff first.

Watch out