Field Record No. 1 documented a governed AI build that survived its executor's death, and it surfaced six operational findings. Every finding was ruled, fixed, and pinned. This record is the re-run: the same plans, the same engine, the same executor model, with the deltas measured. It also caught something new, because a re-verification that finds nothing is suspicious.
Field Record No. 1 documented a governed AI build that survived its own AI process crashing mid-run. That engagement also surfaced six operational findings: things worth fixing. Every one was ruled on, fixed, and locked in with tests. Then we ran the same work again, with the same plans and the same AI model, to measure whether the fixes actually changed outcomes. This record is that re-run.
Nothing about the AI got smarter between these runs. The governance did: better plan authoring, better operator controls, honest failure classification. The improvement is measurable only because both runs kept complete records, and both records are published so you can check the subtraction.
Every improvement process claims to work. This is what one looks like when it has to prove it: problems found on the record, decisions ruled on the record, fixes verified by re-running the work. That loop, findings to rulings to fixes to re-verification, is the difference between a vendor saying "we fixed it" and a vendor showing you the before-and-after records of the same job.
This is the complete execution record of a remediation re-run. During it:
This page is not a simulation. It presents the actual record generated by Atlas Orchestrator, unaltered, and every entry on it opens.
The engine behaved identically in both engagements: every gate fired, every verification was the orchestrator's own. What changed between the runs was the ruled remediation work: plan authoring for headless execution, the cockpit's plan-lineage and supersede flow, honest exit classification, and recovery surfaces. The record measures the difference:
| Record No. 1 (July 8) | This re-run (July 9) | |
|---|---|---|
| Act one, executor sessions | 3 sessions across 2 run rows | 1 session, 1 row |
| Act one, outcome path | crash classification, hand-seeded recovery file through the guidance inbox | uninterrupted; done in 5m 45s |
| Checkpoint session deaths | 2 | 0 |
| Freezing the successor plan | blocked by the cockpit; executed programmatically through the real gates | the supersede dialog, in the cockpit, operator's finger on Freeze |
| Unrequested mid-run work commits | 4 | 0 across nine nodes |
| Act two, wall clock | ~3m (after recovery) | 3m 00s, first try |
| Findings produced | 6 (F-R1–F-R6) | 3 new, dispositioned (see below) |
Governance quality lives in the plan and the cockpit, and both are auditable. Nothing about the AI got smarter between these runs. The difference is remediation, and it is measurable precisely because both runs kept complete records.
Record No. 1's six findings, each discovered by a real operator during real operation. Every one was ruled with its reasoning on the decision log, fixed in a reviewed change with tests pinning the behavior, and re-verified in this run:
| Finding | What it was | The fix, as ruled | Re-verified |
|---|---|---|---|
| F-R1 | Supervised-mode default × headless executor: the session politely died at every checkpoint; recorded as a crash | Plans authored for their run shape (continuous mode); mode disclosed at launch; checkpoint stops now classify paused, never crashed | 1 session/act ✓ |
| F-R2 | A crashed run could not be resumed or relaunched | Relaunch on the run's own view, lineage-linked; resume stays paused-only | not needed ✓ |
| F-R3 | The executor committed per-node on its own judgment | The plan's own text carries the commit discipline; the engine had absorbed the stray commits honestly either way | 0 stray commits ✓ |
| F-R4 | The crashed-run screen was a dead end | Terminal footer: outcome, the engine's own words, next actions | on screen ✓ |
| F-R5 | No way to hand the executor guidance at launch | Operator notes on Launch, Resume, and Relaunch, delivered at the first checkpoint, on the chain | available ✓ |
| F-R6 | The cockpit could not attach, render, or freeze a successor plan | Plan lineage, select-and-render, and a guarded one-way supersede-freeze | used live ✓ |
Each row is one line of the audit log, HMAC-linked to the one before it. Click any entry to see the verbatim record and its hash linkage. Green is the orchestrator verifying; amber is the human acting; rust is a gate refusing, and this run has one.
The successor plan was attached in the cockpit, rendered, validated, and frozen through the supersede flow, the exact flow whose absence was Record No. 1's finding F-R6. At launch the orchestrator verified act one's entire chain, archived it (the honored pause request preserved inside), and narrated the hand-off as its new chain's first entry. It's the first row below. Open it.
Four nodes over the repository act one built: extract a module, prove the old tests still pass, extend coverage, commit once. No interventions, no retries burned, one governed commit.
540508a refactor(run): datefmt v2 governed run work ← act two: one governed commit e33dd38 feat(run): datefmt v1 governed run work ← act one: one governed commit f3cb21f docs: CLAUDE.md operator guide ← the gate's remedy: the bootstrap file, accounted for 973ea8d-era baseline (README only) ← the only ungoverned commit this repo will ever have
A remediation run that reports nothing new should worry you. This one surfaced three findings, each already dispositioned on the decision log: a hash-computation seam whose failure mode had been carrying production honestly the whole time (it labeled what it had rather than guessing what it didn't); a run console that refreshed only on re-entry; and an attention dashboard still using pre-lifecycle definitions of what deserves attention. The first was fixed and pinned the same night; the other two are ruled and in the fix queue. They will be re-verified the same way these six were.
Both chains verify against their keyed genesis: 23 of 23 and 16 of 16 entries. Compare Record No. 1: same plans, 1 crash, 1 recovery, 3 sessions. The delta is the remediation, and both records exist so you can check the subtraction.
This is the loop an audit practice sells: findings stated plainly, rulings with reasoning on the record, fixes reviewed and pinned, and a re-run that measures the difference and hunts for what's next. We run our own tools through it first. No. 3 will be the run where the gates fire on purpose.