What the Bloom established
Reference habitats, causal-selection controls, replay evidence, and the next automated gate.
The Bloom qualification asks a narrow question: can one checked v4 habitat derive two variants from one seed, test both in the world, and use an observed physical result to choose a later trip?
This evaluation records a bounded proposal milestone; it does not claim the complete game or a scientific result.
Qualification result
Eight supplied worlds passed: opposite directions, delayed links, a rotated arrangement, and two crossing worlds whose source route closes and reopens during the confirmation phase. Every world completed 128 ticks. Both builders paid for their rewrites; both children retained their immutable birth provenance. The reference batch used 16 engine executions.
Six controls completed the same horizon and failed Bloom: no edits, equal-direction edits, a preselected winner, a false report relay, no confirmation gate, and a confirmation gate enabled too early. Their receipts remain alongside the references. The sealed qualification used 28 executions within a fixed 128-execution allowance, with source and release-binary identities checked before each call and again at sealing.
The controls separate service from learning. In particular, a preselected child can still perform useful physical work, but it does not earn the Bloom because its choice was not caused by a depot report after both trials. A false relay cannot turn an unsupported claim into selection evidence.
Replay and recovery
The core Bloom tests use 30 measured engine executions: all eight fresh references, prefixes around selection and restoration, a rehashed edit or selector forgery, and a hardwired service control. The CLI tests cover zero-execution case exports, saved prefixes, export/import with pending confirmation, idempotent retries, and cost or receipt tampering; their source estimate is 91 executions and their subprocess telemetry is retained with the run.
The browser bridge has six contract tests. Five synthetic tests passed with 129 assertions; the remaining published-case assertion waits for the recorded public evidence. These tests validate projection identity and visibility. They do not replace Rust replay.
Scope and next gate
This is one fixed eight-cell v4 envelope, one generation, two editable direction loci, and one confirmation. It does not establish arbitrary program synthesis, unbounded evolution, scientific novelty, a P-versus-NP result, multiplayer markets, or the capacity of a larger ecology. The in-world activation ceiling is 16,384 for v4 so a bounded full-body edit can be paid for; v1–v3 retain their 1,024 ceiling.
The next automated gate is an agent study with two bounded ambitions and three candidate slots each. Every candidate must run the four public training worlds under a 64-execution arm allowance; a selected candidate then runs four unchanged transfer worlds. Selection uses checked Bloom outcomes and total modeled work, with every failed candidate retained. After that, measure two fixed workloads and fresh history reads, then compose Bloom with the earlier port commitments before asking people to play.
Modeled work counts in-world operations. It is not wall-clock capacity, token cost, or a proof that one program is optimal. The campaign, complexity notes, and design validation keep those claims separate.
The first bounded agent study now exercises that gate. A keep arm and a frugal arm each admitted two reference-derived candidates, retained one malformed submission, ran all four training worlds, and transferred the selected program across all four unchanged transfer worlds. Both arms completed 24 engine executions and bloomed in transfer; deterministic scoring selected mirror-courier in both because the submitted programs were intentionally equivalent. This establishes that the study protocol, retention rules, and transfer accounting run end to end. It is a harness feasibility result, not evidence that one agent searched better than another. The compact study record and compressed arm ledgers preserve the exact calls and receipts.