Record class: Progenitor practice record. Presentation revision 3: terminology and claim-boundary corrections applied per review; revisions 1 and 2 preserved byte-exact in this archive; underlying record bytes unchanged. Historical classification: progenitor practice record (remediation), produced by Atlas's pre-formal governed execution machinery; not a v1.3 core bundle, no formal GAD conformance result, and it predates the bound-probe floor. Historical evidence label: server_verified; current interpretation: engine capture of the check, credited at proxy strength or below under today's registry.
Atlas North Institute · field record No. 2 · July 9, 2026

Six findings. Six fixes.
Then we ran it again.

Field Record No. 1 documented a governed AI build that survived its executor's death, and it surfaced six operational findings. Every finding was ruled, fixed, and pinned. This record is the re-run: the same plans, the same engine, the same executor model, with the deltas measured. It also caught something new, because a re-verification that finds nothing is suspicious.

guided tour about 4 minutes · full evidence review about 20 · or scroll freely, every entry opens
plan v1 ad004fd5e54b…
plan v2 b20ae029d3fe…
engine Atlas Orchestrator 1.3.0
cockpit Waypoint (post-remediation)
Executive briefing · about 3 minutes · every statement below is drawn from the technical record on this page

The setup

Field Record No. 1 documented a governed AI build that survived its own AI process crashing mid-run. That engagement also surfaced six operational findings: things worth fixing. Every one was ruled on, fixed, and locked in with tests. Then we ran the same work again, with the same plans and the same AI model, to measure whether the fixes actually changed outcomes. This record is that re-run.

What happened in this engagement

  1. The build that previously took three AI sessions and a crash recovery completed in one uninterrupted session, 5 minutes 45 seconds.
  2. A safety gate fired for real: the system flagged a file the AI touched that was not accounted for in the plan. The AI corrected it, and verification passed. 23 seconds from refusal to verified fix, all on the record.
  3. The follow-up plan was launched through the operator's console, with a human's finger on the button, using controls that did not exist before the first record's findings.
  4. The hand-off between plans was verified and archived by the system itself.
  5. The re-run surfaced three new findings, each logged and dispositioned, because a re-check that finds nothing new should worry you.

The measured improvement

3 → 1
AI sessions needed for the same build, first record versus this one
2 → 0
session deaths during the build
1 gate firing
a real refusal, remediated and re-verified in 23 seconds
39
HMAC-linked log entries (verified end to end within the engagement trust model); both chains verify end to end

Nothing about the AI got smarter between these runs. The governance did: better plan authoring, better operator controls, honest failure classification. The improvement is measurable only because both runs kept complete records, and both records are published so you can check the subtraction.

What this means for your business

Every improvement process claims to work. This is what one looks like when it has to prove it: problems found on the record, decisions ruled on the record, fixes verified by re-running the work. That loop, findings to rulings to fixes to re-verification, is the difference between a vendor saying "we fixed it" and a vendor showing you the before-and-after records of the same job.

What this record proves, and what it does not

In plain terms: this record is HMAC-linked and verified end to end within the engagement trust model; it carries no independently verifiable public time anchor, and altering the stored record without breaking its chain would require the relevant key and custody assumptions to fail. The fixes' effect is shown by two complete records of the same work, checkable by subtraction rather than taken on our word. It does not prove the software built is good, secure, or compliant. Those are human judgments, made separately. A record that claimed more would be worth less.
every log entry opens, including the gate firing · about 20 minutes
In one minute

What you're about to see

This is the complete execution record of a remediation re-run. During it:

This page is not a simulation. It presents the actual record generated by Atlas Orchestrator, unaltered, and every entry on it opens.

What this does and doesn't prove. Checks labeled server_verified were executed by the orchestrator itself, never taken on the AI's word. The chain makes later substitution detectable within the stated trust model: it establishes the identity, sequence, and internal linkage of what was recorded, not that every recorded event physically occurred as described. The before-and-after records show different observed outcomes under the two executions. Neither is a claim that the software built is good. That distinction is the product.
The measured deltas

Same plans, same engine, same executor model

The engine behaved identically in both engagements: every gate fired, every verification was the orchestrator's own. What changed between the runs was the ruled remediation work: plan authoring for headless execution, the cockpit's plan-lineage and supersede flow, honest exit classification, and recovery surfaces. The record measures the difference:

Record No. 1 (July 8)This re-run (July 9)
Act one, executor sessions3 sessions across 2 run rows1 session, 1 row
Act one, outcome pathcrash classification, hand-seeded recovery file through the guidance inboxuninterrupted; done in 5m 45s
Checkpoint session deaths20
Freezing the successor planblocked by the cockpit; executed programmatically through the real gatesthe supersede dialog, in the cockpit, operator's finger on Freeze
Unrequested mid-run work commits40 across nine nodes
Act two, wall clock~3m (after recovery)3m 00s, first try
Findings produced6 (F-R1–F-R6)3 new, dispositioned (see below)

Governance quality lives in the plan and the cockpit, and both are auditable. Nothing about the AI got smarter between these runs. The difference is remediation, and it is measurable precisely because both runs kept complete records.

The audit loop

Findings → rulings → fixes → re-verification

Record No. 1's six findings, each discovered by a real operator during real operation. Every one was ruled with its reasoning on the decision log, fixed in a reviewed change with tests pinning the behavior, and re-verified in this run:

FindingWhat it wasThe fix, as ruledRe-verified
F-R1Supervised-mode default × headless executor: the session politely died at every checkpoint; recorded as a crashPlans authored for their run shape (continuous mode); mode disclosed at launch; checkpoint stops now classify paused, never crashed1 session/act ✓
F-R2A crashed run could not be resumed or relaunchedRelaunch on the run's own view, lineage-linked; resume stays paused-onlynot needed ✓
F-R3The executor committed per-node on its own judgmentThe plan's own text carries the commit discipline; the engine had absorbed the stray commits honestly either way0 stray commits ✓
F-R4The crashed-run screen was a dead endTerminal footer: outcome, the engine's own words, next actionson screen ✓
F-R5No way to hand the executor guidance at launchOperator notes on Launch, Resume, and Relaunch, delivered at the first checkpoint, on the chainavailable ✓
F-R6The cockpit could not attach, render, or freeze a successor planPlan lineage, select-and-render, and a guarded one-way supersede-freezeused live ✓
Act one · plan ad004fd5 · 23 chained entries · one session, 5m 45s

The chain, with a gate firing in it

Each row is one line of the audit log, HMAC-linked to the one before it. Click any entry to see the verbatim record and its hash linkage. Green is the orchestrator verifying; amber is the human acting; rust is a gate refusing, and this run has one.

The hand-off

Succession, through the cockpit this time

The successor plan was attached in the cockpit, rendered, validated, and frozen through the supersede flow, the exact flow whose absence was Record No. 1's finding F-R6. At launch the orchestrator verified act one's entire chain, archived it (the honored pause request preserved inside), and narrated the hand-off as its new chain's first entry. It's the first row below. Open it.

Act two · plan b20ae029 · 16 chained entries · one session, 3m 00s

The refactor, clean on the first try

Four nodes over the repository act one built: extract a module, prove the old tests still pass, extend coverage, commit once. No interventions, no retries burned, one governed commit.

the repository's final history · git log --oneline, verbatimworking tree clean
540508a refactor(run): datefmt v2 governed run work  ← act two: one governed commit
e33dd38 feat(run): datefmt v1 governed run work      ← act one: one governed commit
f3cb21f docs: CLAUDE.md operator guide               ← the gate's remedy: the bootstrap file, accounted for
973ea8d-era baseline (README only)                   ← the only ungoverned commit this repo will ever have
The honest asterisk

The re-run found three new things

A remediation run that reports nothing new should worry you. This one surfaced three findings, each already dispositioned on the decision log: a hash-computation seam whose failure mode had been carrying production honestly the whole time (it labeled what it had rather than guessing what it didn't); a run console that refreshed only on re-entry; and an attention dashboard still using pre-lifecycle definitions of what deserves attention. The first was fixed and pinned the same night; the other two are ruled and in the fix queue. They will be re-verified the same way these six were.

Run facts

The numbers, all real

2
governed plans, one repository
9
nodes of AI-written code
39
chained entries, every one on this page
9
server-owned verification passes
1
verification failure: a gate firing
23s
from gate refusal to remediated pass
1
operator pause, honored at a checkpoint
1
verified succession
0
crashes · halts · waivers · healings

Both chains verify against their keyed genesis: 23 of 23 and 16 of 16 entries. Compare Record No. 1: same plans, 1 crash, 1 recovery, 3 sessions. The delta is the remediation, and both records exist so you can check the subtraction.

No. 1 proved the record survives failure. No. 2 proves the findings get fixed.

This is the loop an audit practice sells: findings stated plainly, rulings with reasoning on the record, fixes reviewed and pinned, and a re-run that measures the difference and hunts for what's next. We run our own tools through it first. No. 3 will be the run where the gates fire on purpose.

Atlas North Institute · AI governance & assurance · re-run of 2026-07-09 · engine Atlas Orchestrator 1.3.0 · cockpit Waypoint · identities ad004fd5e54b / b20ae029d3fe
Verification panel (presentation rev 3)
Presentation: field-record-2.html · rev 3
Underlying run project: project-4-1783560942369 (cited by the record)
Artifact location: REQUIRED: predates the supplied Waypoint runs corpus; locate in the original installation archive
Chain heads on page: 2a2ad6331b71782eed9e686ab45b22a45747d41e9067c5eabfa991a82c74642d · 610fa76bc1e96f24cc5b5bf8662364a130c62f09bed4b9a8016314e7c0290ef2 (as printed)
Protocol: pre-formal Atlas machinery
Anchor: none claimed