Record class: Trial 0 record (formal GAD protocol record). Presentation revision 3: terminology and claim-boundary corrections applied per review; revisions 1 and 2 preserved byte-exact in this archive; underlying record bytes unchanged. Post-publication review status: fifth cross-model review accepted formal findings affecting predicate domains, claim typing, policy completeness, and schema normalization. No accepted finding has yet been shown to invalidate this identified Trial 0 result. Corpus-impact replay under the named evaluator, policy, and bundle root is pending and will be published as an identified record.
Read this as
Atlas North Institute · field record No. 6 · July 16, 2026

The referee said no to its own makers. Then it said yes, and it meant something.

On July 16 the production stack tried to produce the first live production-stack bundle evaluated VALID under the identified v1.3-track evaluator — and the separately executable evaluator rejected its makers' work twice, by named rule and entry index, before accepting the third attempt: VALID_RECORD=true, all nine conditions, independently timestamped by an RFC 3161 authority, establishing the accepted record's existence no later than the token time, on a second machine, cold. A tested one-byte mutation, byte 4000 of the witness, caused conditions R1 and R2 to fail and the evaluator to return INVALID; additional invalid vectors exercise the other stated conditions. This is the record of the refusals and of the yes they purchased.

What happened

Every prior record in this archive was verified by machinery that shared a roof with the tools being verified. Trial 0's second movement demanded more: a live run of the production stack — the desktop supervisor driving the orchestration engine driving a real AI executor — whose exported record a separate program, gad-evaluate, would judge on another computer with no access to the producer's software, state, or secrets.

The final evaluation, computed cold on a Linux machine from a record produced on Windows:

VALID_RECORD=true · OUTCOME=COMPLETED · 9 of 9 conditions · referee exit 0. Record root bf7ab081…f83abb, 29 witness entries: the frozen plan first, two executor sessions, ten probe artifacts, ten verdicts each recorded outside any session over a captured delta, and a terminal that waited for the last session record to close before sealing.
Anchored. A live RFC 3161 timestamp from a public authority, its message imprint equal to the record's root. The record now makes a claim about when that no one who controls the machines can quietly move.
One flipped byte → INVALID. Byte 4000 of the witness, flipped, re-evaluated: conditions R1 and R2 fail, exit 1. The clean verdict is falsifiable, which is why it is worth something.
Zero key material. The exported bundle swept by derivation — 9 files, 359 candidates, 0 findings — using the checker the specification made normative the same day.

The part worth reading

The yes is the headline. The refusals are the proof. Three earlier attempts that day produced, in order: no record at all (the plans never armed the protocol layer — four governed runs, chained and complete, and structurally silent); a record rejected seven times over on one rule (verdicts recorded inside later sessions, after their nodes had been re-dispatched — the evaluator's transition table demands verification happen between sessions, in nobody's session, and the product had no such caller); and a record invalidated by its own operator's mouse click (a relaunch after the terminal sealed — legal-looking in the UI, illegal in the record, and correctly fatal).

Each rejection changed something permanent. The first produced a three-layer identity rule: before any record-producing run, prove the plan, the engine, and the supervisor are the ones you think they are. The second forced the product to grow the between-sessions verifier the specification always implied — the evaluator, in effect, filed a feature request and made it non-negotiable. The third was ruled a product defect on the spot, in the operator's words: a completed run must be inert. The guard is queued; the click that found it is preserved in this archive, attributed to the advisor who recommended it.

A verification regime you built cannot impress you by approving you. It can only impress you by refusing you until you deserve it — and showing its arithmetic both times.

The full technical record — the three rejections with the referee's verbatim rulings, the annotated 29-entry witness, and the commands to re-verify everything — is below.

Three rejections, verbatim

Rejection 0 · the record that never existed. Four succession runs executed under classic governance — approval gates, server verification, unbroken hash chains — but their plans omitted safety.protocol_mode: "v0_6", so the witness layer never opened. The export said so instead of faking one: a stated absence, not a fabricated record. Register entry REG-53; doctrine: the plan carries the protocol, the engine is the resolved index, the supervisor is the build that drives the records — verify all three before a run whose product is a record.
Rejection 1 · run 15. R3: FAIL — verdict for "m2-01-referee-whole" without a completed attempt (a FAILED executor outcome routes to retry/trip, never to verification) — seven counts. Root cause, verified against the evaluator's source: the transition table resets a node's completed attempt at every dispatch naming it, and the supervisor commissions each session over the whole remaining plan; so a verdict recorded in a later session always trails a fresh dispatch and is illegal. Verification belongs between sessions. The product had no such caller. Register entry REG-54; the between-sessions verifier was built, tested against the evaluator's own trace-checker, and its probe reproduces run 15's exact defect and watches the referee name it at the right index.
Rejection 2 · run 18. R3: FAIL — trace index 18: dispatch after the terminal entry (SEALING) and R5: FAIL — worker-authored proxy credited with no probe bound, two counts. The first was the operator relaunching a finished run at the advisor's suggestion — the witness had sealed; the click appended to a sealed chain. Ruled a defect in the product, not the operator: nothing should accept a dispatch over a terminal-carrying witness. Register entries REG-55 and REG-56. The second was two remaining unprobed checks in the plan; the final plan binds a probe to every provable check — ten pairs, twenty commands, all fired in the pre-freeze matrix.

The valid record: 29 entries, annotated

  0  freeze                      the frozen plan's hash and canonical bytes: recorded precedence
  1  dispatch [m2-01, m2-02]     session 1 commissioned over the node set, before the executor exists
  2  executor_exit    completed  the session ends cleanly; node null: a multi-node session names no single owner
  3  execution_result            the engine's capture of what the session left
  4-10  probe (7)                the negative evidence, recorded as artifacts: each check's planted-defect probe
 11-17  verdict m2-01 pass (7)   recorded BETWEEN sessions by the supervisor's verifier, over the captured delta
 18  dispatch [m2-01, m2-02]     session 2; m2-01's verdicts all precede this line, which is the entire point
 19  executor_exit    completed
 20  execution_result            the second session's capture: one authored file, inside its fence
 21-23  probe (3)
 24-26  verdict m2-02 pass (3)   the final node's verdicts, again in nobody's session
 27  plan_completed              the terminal, DEFERRED until the last session record closed
 28  checkpoint                  post-terminal bookkeeping, one of exactly two entry types allowed after a seal

The content the record attests is itself the project's governance: node one re-proved the entire sealed succession lineage of field record No. 5 inside the witness; node two authored the Movement 2 note whose every figure was computed live — the lineage verifier's own final line quoted verbatim, the v0.9 digest read from the signed succession record, never retyped.

Identity: what ran, and what a stranger checks

record root (root_core) bf7ab08163ec442caaac6e997e7d72b89772f2a3b0d89ed36fc4e44b8bf83abb
anchor RFC 3161 token · https://freetsa.org/tsr · message imprint = root_core
engine key shipped in verification-inputs.json as public SPKI material; verifies every engine-signed entry
judgment gad-evaluate <bundle>/envelope --policy gad4-default.json --keys <engine pubkey> --json → exit 0

The evaluator needs the bundle, the policy, and public keys. It does not need the producing machine, the supervisor, the engine, this institute, or anyone's goodwill.

The boundary this record refuses to blur

VALID is not DEFENSIBLE. Under the strictest policy overlay, this record computes DEFENSIBLE=false: the overlay wants an authority-signed anchor, and the record carries a timestamp authority's token — a claim about when, not about who stands behind it. That gap is an open ruling in the project's register, stated here rather than rounded up. And per the specification's own header, every result in this record is a Proposition or Claim until independent formal and cryptographic review completes. One more honest edge: both of this record's sessions were multi-node; the single-node session tier is demonstrated in the engine's conformance harness, not yet in a live record.

What this record proves, and what it does not

It proves: the full production stack can produce a record that a separately executable evaluator accepts, on a different machine and operating system, for stated reasons, with tamper-evidence demonstrated by construction and by experiment; that the record's timeline is anchored to a public authority's clock; and that the same evaluator rejects this same stack's work when the choreography is wrong — by rule name and entry index, twice, on the same day.

It does not prove: that the work inside the record is good (the record proves the checks ran and passed; the checks' worth is a human judgment); that the stack is conformant in general: the named v1.3-track evaluator returned VALID for this exact bundle under the identified inputs, a historical computation that remains attributable and reproducible, and broader tooling conformance is not inferred from it; or that the evaluator itself is correct (that is Phase 5's question, for reviewers who owe this project nothing).

Verification panel (presentation rev 3)
Presentation: field-record-6-movement-2.html · rev 3
Bundle root on page: bf7ab08163ec442caaac6e997e7d72b89772f2a3b0d89ed36fc4e44b8bf83abb (the accepted record)
Artifact location: REQUIRED: Trial 0 bundles live in the gad-protocol conformance tree, not the supplied Waypoint runs corpus
Evaluator / policy: gad-evaluate (version and digest REQUIRED at publication) / gad4-default (digest REQUIRED)
Anchor: RFC 3161 token over the accepted record; token digest REQUIRED at publication
DEFENSIBLE: false under the named policy, as the page states