In 63 minutes, a governed run built gad-evaluate: the separately executable evaluator that verifies Governed AI Development records on a stranger's machine. Fifteen nodes, zero retries, zero halts, 61 chained entries. The run that ended the era of "trust our tooling" was itself governed by the tooling it retires. Every verification, class label, and warning is preserved and inspectable on this page.
On the night of July 13, the reference machinery of Governed AI Development executed a frozen fifteen-node plan and built the discipline's missing piece: a small, separately executable evaluator that answers three questions about any governed record, on any computer, with no access to the producer's software or secrets. Is the record intact and internally lawful? How did the run actually end? Was the contract satisfied? It signs its answers. The same night, on a second machine running a different operating system, a cold rebuild passed all seventeen exit checks, and the evaluator rejected nine deliberately broken records, each for the correct stated reason, caught a single flipped byte, and refused to emit an unsigned attestation, because, in its own words, an unsigned attestation is not an attestation.
A vendor showing you a flawless run has shown you nothing: flawless is what fabrication looks like too. This record's clean report borrows its credibility from its siblings. The honesty box above never disappears from these records, and in record No. 1 it reported a mid-run executor crash; in No. 2, the failures being remediated. Same machinery, same box, same refusal to hide. When that machinery reports a clean run, the report means something, because you have seen what it says when things go wrong.
And this plan was not trusted on sight. Before launch, the engine's own strict validator rejected the first draft: the genesis node did not prove its baseline commit, and every checker carried a stronger evidence label than an executor-authored tool can honestly hold. The plan that ran was the corrected one, every checker demoted to its true class and required to demonstrate, against a committed planted defect, that it can fail. Fifteen green nodes, each one green only after its checker demonstrated a failure under its planted-defect probe. see the run's first warning, honestly recorded →
What follows is analysis, not chained evidence: the record proves the run's duration and interaction pattern; it cannot prove a counterfactual that never executed. The governed run cost one launch and fifteen brief supervised checkpoints, because the frozen plan carried the entire specification and the executor never had to be told anything twice. The same scope executed ungoverned, estimated from this operator's own prior build rhythm, runs seventy to one hundred ten authored prompts across several days, most of them specification retyping and correction cycles. The largest item in that estimate is exactly what this record shows never happened: zero retries means the correction loop, which is most of what ungoverned prompting is, never ran once. Roughly a four-to-six-times specification efficiency, and the ungoverned version produces this record at no price: never.
| Frozen plan (authored identity) | 0683ba78ad295e08fa6077d585951044100606ebf8360ab8193ef208258920b0 |
| Master spec hash | 71a0c23460ea2da1faef5c9dcc5ee89c1dec6477ecc8f2061425ef1f809eed4a |
| Genesis hash (chain root) | dc9e7bddb9e3bfa191dd16a526ee0e8f807bfbd95bb1c8d2e8633eaede6eb852 |
| Engine · cockpit | Atlas Orchestrator 1.5.0 · Waypoint 1.1.0 |
| Nodes · retry cap · mode | 15 · 3 · supervised, scope enforced, server verification default |
| Repository at start | none: genesis initialized the repository; the one pre-seeded entry (docs/) was warned about on the chain, entry 1 |
Durations are claim-to-verification, computed from the entries below. Every check was run by the engine against executor-authored tooling, so every row carries the honest class for that arrangement, server_observed_proxy, and every checker holds a negative probe: a committed planted defect it must detect, or its passes do not count.
| # | node | title | duration | checks | evidence class |
|---|---|---|---|---|---|
| 01 | n01-scaffold | Genesis: repo scaffold, script surface, standing rules | 8m 20s | 2 | server_observed_proxy |
| 02 | n02-register-mirror | Register verification: seeded REG-5 closure and REG-6 through REG-9 | 1m 21s | 1 | server_observed_proxy |
| 03 | n03-schema-signed-witness | Schemas: signed-object and witness-entry | 10m 28s | 2 | server_observed_proxy |
| 04 | n04-schema-plan | Schema: plan (P) | 1m 05s | 2 | server_observed_proxy |
| 05 | n05-schema-bundle-envelope | Schemas: core-bundle and envelope (REG-7, REG-8 form) | 1m 24s | 2 | server_observed_proxy |
| 06 | n06-schema-operator-policy-attestation | Schemas: operator, policy, attestation | 1m 49s | 2 | server_observed_proxy |
| 07 | n07-canonicalization | Canonicalization: RFC 8785 base, domain tags, vectors | 2m 57s | 1 | server_observed_proxy |
| 08 | n08-transition-table | Transition table for L(C) and invalid traces (REG-1, REG-3, RUL-2/3/4) | 5m 31s | 2 | server_observed_proxy |
| 09 | n09-registries | Predicate and evidence-strength registries | 1m 53s | 1 | server_observed_proxy |
| 10 | n10-default-policy | Default GAD-4 trust policy | 1m 04s | 1 | server_observed_proxy |
| 11 | n11-fixtures-valid | Valid fixture envelopes: COMPLETED, HALTED, INCOMPLETE | 3m 14s | 2 | server_observed_proxy |
| 12 | n12-fixtures-invalid | Invalid fixtures: one per R-condition | 2m 20s | 1 | server_observed_proxy |
| 13 | n13-evaluator-core | gad-evaluate core: R1 through R9, OUTCOME, CONTRACT_SATISFIED | 5m 38s | 2 | server_observed_proxy |
| 14 | n14-evaluator-cli-q | CLI and attestation Q | 2m 24s | 1 | server_observed_proxy |
| 15 | n15-phase-gate | Phase 1 exit gate: the full acceptance sweep | 1m 27s | 1 | server_observed_proxy |
It proves the plan was frozen and hash-addressed before work began; that every completion claim was checked by machinery outside the worker's session with its evidence class stated per node; that every checker demonstrated the capacity to fail before its passes counted; that the history is chained end to end under the engagement trust model; and that the resulting code passed its full exit gate twice, the second time cold, on a machine its builder never touched. It does not prove semantic correctness against the specification, the efficiency estimate above, or this record's own conformance to the specification its contents implement. Those verdicts belong to the adversarial vectors, the second evaluator, and Trial 0, all chartered, none skipped.
All 61 entries, seq 0 through 60, exactly as written at the moment. Click any entry for its raw form and its link to the next.