What earned the answer

Frozen agent comparisons, physical signal provenance, restored journeys, and the measured cost of the complete local adventure.

The First Answer tests whether a familiar courier, useful construction, continued service, and a returned signal can compose into one saved adventure. The larger campaign remains a proposal. The recorded journey exposes the actual Rust frames. The walkthrough reaches the same local ending through the agent-facing CLI.

One crew, one deadline

The world has five initial cells, two supplied child blueprints, two material tokens, four total links, and six physical sparks. Both service beacons must receive three matching sparks and remain supplied through 128 ticks. All work fits inside the original 40,000-unit allowance; there is no refit, external spawn, or fuel purchase during the save.

In the supplied reference, the courier and Keeper retain their earlier programs. The builder constructs Keeper first, then a reply cell. The relay forwards depot reports into that cell, which returns a report to a waiting receiver. Construction, service, and contact each have their own observable result.

The ending joins several facts checked by Rust: the courier physically deposited the final spark; the born Keeper routed it to service; the born reply cell sent a report carrying that spark's identity and bit; the receiver consumed the report after the physical delivery; and service survived the complete horizon. The material used by both children must come from the declared cache.

This is a report originating at a depot, joined with independently checked service. The beacon does not emit an acknowledgment. The ending text is authored. The experiment demonstrates a finite communication and construction task, without establishing language, consciousness, or an autonomous civilization.

References and causal controls

All eight reference worlds qualified on their first attempts. Keeper activated at tick 17 and the reply cell at tick 38. Their supplied bodies contained 444 and 539 canonical bytes respectively. Final matching contact occurred between ticks 94 and 111, inside the 128-tick horizon. The authored ending remained unavailable until that horizon finished.

Four additional qualification controls removed the second material token, disabled the return link, replaced the reply with a constant, or idled the courier. The first three preserved service but failed to earn the answer. The idle courier failed both. A correct zero is therefore insufficient: the successful zero report carries the actual spark identity, while a default or constant zero does not.

The ending deliberately requires a verifiable exchange. These controls test that contract; they do not establish that communication always improves service or that a more complex organism is inherently better.

Qualification preserved all twelve input/receipt pairs and used exactly 24 engine executions: one run and one fresh verification per attempt. The journey projection used those already verified traces without another execution.

Two player ambitions

Two independent developer agents each received three submitted candidate slots. Every candidate faced four training worlds. Selection required answering all four, then minimized total modeled work, canonical program bytes, and submission order. Each selected program bundle was frozen before four public transfer cases.

“Keep my courier” preserves the earlier courier exactly while permitting builder and reply-policy edits. “Spend less work” additionally permits courier edits. All other bodies, links, supplies, initial state, and schedules remain fixed. The broader ambition may choose the same courier; different winners are not manufactured through scoring bonuses.

The public reference is admissible. Failed submissions and executions stay in the record. Agent reasoning and development context are available to the authors, so these are developer diagnostics, not blind agent rankings. Their external token and monetary costs are not observed.

Both agents completed all twelve training journeys and all four unchanged transfer journeys. Every submitted candidate was valid and successful; no slots were discarded. Each arm used 32 actual engine executions against its 64-execution allowance.

Ambition and submitted candidateTraining answersTraining workCanonical program bytes
Keep: live-echo4 / 488,0455,409
Keep: direct-echo4 / 486,4525,312
Keep: rested-builder, selected4 / 483,1245,253
Frugal: quiet-workshop4 / 483,5935,186
Frugal: late-compass-setup4 / 480,1855,050
Frugal: compact-detours, selected4 / 477,3674,356

The Keep selection remembers when construction has finished and preserves the exact inherited courier. Its four transfer cases used 80,834 work, for 163,958 across all eight worlds. The Frugal selection used 75,093 transfer work, for 152,460 across all eight.

The Frugal agent kept its builder and reply programs fixed across its three candidates. Courier edits reduced training work by 6,226 units, about 7.4%, while preserving every observed courier action and position. These savings came from a smaller description and fewer interpreter checks, rather than a faster physical journey. Changing the courier therefore had measurable value under this allowance. This specialized courier is qualified on these eight worlds; it has not inherited the earlier courier's separate 44-world navigation result.

The shared comparisons also include initially installed children and a reply policy that forwards its current inbox directly. Initially installed children deliberately change the starting resources; this measures the cost of requiring assembly, rather than offering a legal shortcut within the construction task. Direct forwarding uses the ordinary interpreter and can legitimately outperform a more elaborate remembered reply.

Across all eight fixed worlds, direct forwarding answered every case and used 4,936 fewer modeled work units than the reference, a reduction of about 2.8%. The simpler program therefore wins this comparison.

Shared comparisonService successesConstructed journey answersTotal modeled work
Supplied reference8 / 88 / 8178,736
Direct forwarding8 / 88 / 8173,800
Initially installed children8 / 80 / 8148,326

The larger control set tested six removals or substitutions on two opposite-bit cases each. Removing stock, activation, or courier behavior failed both service and the ending. Idling the reply cell, disabling its return link, or latching an old report preserved service in both cases while failing the ending. These twelve controls all ran to completion; their failed answers remain valid, replayable results.

A save that survives interruption

Four trajectories used all eight permitted advances, pausing at ticks 1, 2, 8, 17, 18, 38, a predeclared reference-contact checkpoint, and 128. Actual contact milestones are reported separately: a selected crew can reply earlier, while the old-report control never earns matching contact. Each trajectory was exported and restored at ticks 8 and 38. The completed results and journey grades matched the corresponding cold runs exactly. Each trajectory consumed 304 actual engine executions because commands freshly checked the retained history.

Twelve integrity probes attempted altered or mismatched evidence. All were rejected, consuming 21 engine executions in total. The complete frozen study retained 68 cold journeys and used 1,373 actual engine executions against its 2,048-execution allowance. This includes both agents, shared comparisons, saved replay, and rejected probes; the separate qualification and CLI calibration used 378 of 384 executions.

Measured cost of this adventure

Capacity was measured on an Apple M5 Max running macOS, using the pinned Rust 1.97.1 release build. Each of two reference worlds received one warmup and thirty recorded samples. A sample creates a receipt, freshly verifies and grades it, and serializes it through a fixed reusable 8 MiB buffer. These operations used 124 engine executions in total.

Reference worldRun, verify, grade, serialize p50 / p95Including extra identity checks p50 / p95Receipt bytes
answer-one14.23 / 14.97 ms18.56 / 19.19 ms678,345
answer-crossing13.41 / 13.95 ms17.54 / 18.34 ms682,215

The benchmark process peaked at 6.42 MiB resident memory. Two separate completed-history reads used 17 fresh engine executions each: 66.38 ms with 13.67 MiB peak memory for the reference, and 59.84 ms with 12.70 MiB for the selected familiar courier. These are single observed reads, not latency percentiles. The complete capacity capture used 158 of its 256-execution allowance.

The four finished save bundles occupy 1,409,825–1,656,343 bytes each. Both measured history reads and the benchmark passed the declared 256 MiB memory, 8 MiB receipt, and 16 MiB completed-bundle thresholds. The runtime's separate 64 MiB import limit is not a demonstrated operating size.

Repeated history verification costs more than advancing the tiny world once. That is acceptable for this local adventure and an explicit constraint for longer histories. These measurements qualify seven potential cells over 128 ticks; they do not project a price for thousands of active organisms, larger worlds, a complete campaign, or hosted multiplayer. Model tokens, hosted bandwidth, and monetary costs remain unmeasured.

Inspect and reproduce the evidence

The frozen protocol, complete study ledger, and capacity samples retain the declared allowances, exact inputs, candidate selection, process metrics, failures, and artifact identities. The recorded journey offers four receipts and their save bundles, including an unsuccessful old-report control.

After installing the pinned toolchain and dependencies, run bun run check from the repository root. Its First Answer gate freshly verifies the retained receipts, imports and checks saved prefixes, repeats the declared integrity probes, and validates the public projections. Every admission run has its own 1,024-execution cap and prints its actual engine count, including on failure. Admission and later integration checks are additional verification work, separate from the frozen search and capacity allowances.

What this closes, and what remains

The local objective supplies a beginning, physical construction, continued responsibility for the crew, and an earned ending. Saved prefixes preserve unfinished bodies, reports, fuel, and ancestry. That makes a continuous construction-to-contact adventure executable rather than leaving its final event solely in prose.

The complete Long Trail still has additional contracts:

Campaign dependencyRemaining evidence
Programmable ark controlThe subsequent ark diagnostic checks one addition and two service plans in a fixed habitat; repeated regulation and integration with construction remain
Connected settlementsPhysical commitments under declared message loss, delays, and duplicates, with single custody and no duplicate credit
Endogenous searchAn in-world process generates a changed candidate, pays for every evaluation, and selects a result that passes frozen confirmation
Complete campaignA fresh save traverses every chapter transition with prior creations doing useful work, followed by capacity qualification for that actual journey

Those gates remain before human playtesting. Reachability and reproducibility cannot establish attachment, pacing, or enjoyment. Improvements on these supplied worlds are computational results about these worlds; scientific novelty and claims about P versus NP require separate research questions and evidence.