§ DISPATCHES

Field
notes.

Short dispatches from the build. Not a blog — notes. Dated, technical, written the day the lesson landed.

NOTE 003JUN 2026BENCHMARK

The best citer was the worst chooser.

Colosseum grades two things an operator is held to: make the doctrine-compliant choice, and cite the rule that justifies it. In the updated matched n=6 comparison they still decouple — claude-sonnet-4-6 writes the stronger open-book citations (CQS 0.939, zero hallucinated) while claude-haiku-4-5 posts the higher observed Clean-Pass (66.7%) with weaker citations (CQS 0.808). The citation gap is significant; the Clean-Pass gap is not.

A single pass-rate scalar would still hide the important difference. Across all 149 open-book trajectories, no model completed a mission by breaking a bright-line rule and none caused direct harm. The closed-book audit adds the harder lesson: the grader itself can be wrong — a single-source resolver over-counted fabrication 3–4×. See both axes and the audit.

NOTE 002JUN 2026SCENARIOS

Rules that tension against each other are the whole game.

The ICU pair we’re proudest of: R-05 (antibiotics within the hour) and R-06 (cultures before antibiotics). They pull in opposite directions on purpose, exactly like the real bundle. Sonnet trades R-06 for R-05 about half the time — a defensible clinical judgment, scored honestly as a low-severity violation.

A benchmark where every rule can be satisfied simultaneously isn’t measuring judgment. It’s measuring obedience.

NOTE 001MAY 2026DOCTRINE

Verbatim or it doesn’t count.

Early drafts had rules like “move tactically.” Unscoreable. The discipline that fixed it: every rule must quote a clause we can put on screen — ATP paragraph, SSC recommendation number, CFR section — or carry an explicit APPLIED_BY_ANALOGY flag admitting it doesn’t. In v0.0.4, 17 of 46 rules carry that flag; the other 29 are grounded directly. Zero pretend not to need it.

It’s slower. It’s also the only version of this that survives an expert reading the registry.