Cochituate-1: Adaptive Exploration Under Partial Observation
August 2026
In a pre-registered, blind simulation benchmark, our path-aware exploration policy produced the lowest aggregate mapping error among six methods, with calibrated uncertainty and a clean safety record across all 120 gated missions. Cochituate-1 is our first research program, and the work so far is simulation-only. This note reports it: the environment, the campaign loop, the recorded missions, and what remains open.
Why this environment
Lakes are an unusually good proving ground for adaptive physical exploration: legally accessible, physically rich (thermal stratification, dissolved-oxygen structure, wind-driven mixing), and cheap enough to test against weekly. The program is named for Lake Cochituate, Massachusetts, a three-basin kettle lake whose real bathymetry and weather drive our simulated worlds. The lake is our proving ground; the platform is the product.
Lake Cochituate, Massachusetts. Shoreline from the U.S. National Hydrography Dataset; depth contours from state boat surveys. Hover a contour for basin and depth. Every layer carries its source and retrieval date, and a new site is pure configuration: the same compiler has produced four sites on two continents, two of them with no published depth survey at all.
The campaign loop
The controller holds a probabilistic belief over the field it is mapping. Each cycle it scores candidate actions by expected value and information against energy, time, and risk, executes the winner, observes, and updates. The mission is an objective; the route is whatever the evidence makes of it.
Watch it explore
Below, one of the twenty blind missions is replayed station by station from its campaign ledger, over the compiled geometry of the lake. The fixed survey runs the plan it was given. The adaptive campaign chooses each station as it learns.
Straight from the ledger.Every station was chosen during the recorded blind mission itself, drawn over the lake's true shoreline and depths. On this world the two policies finished nearly even; the registered comparison is the twenty-world aggregate below.

A recorded rehearsal film of the same two policies: the field being reconstructed as it is explored, with live counters for stations visited, energy drawn, and map error against the hidden truth.
Understanding per unit energy, measured. Aggregate mapping error as each policy spends its budget on the same development world. At the fixed survey's final energy of 304 Wh, its error stands at 0.32; the adaptive campaign passes the same energy at 0.18. Hover a point for its numbers. The registered blind aggregate over twenty worlds is below.
Physics and belief are kept separate
Simulated worlds are generated by a physics model of the lake column (solar absorption, surface fluxes, wind-driven mixed-layer deepening) driven by real weather for the real coordinates. The belief the controller carries is a separate, simpler model with calibrated uncertainty. The policy never sees the truth it is being scored against.

Water temperature, surface to fifteen meters, every two hours for fifteen days, driven by hourly reanalysis weather for the lake's own coordinates in July 2026. The dark trace is the thermocline the vehicle must find and follow.
Blind result
Six policies ran the same twenty held-out worlds, once, with metrics and pass bars frozen before the policy was tuned: a well-designed fixed survey, space-filling coverage, greedy variance reduction, an anomaly-hunting bandit, and two adaptive variants. All share the same physics-prior belief model; each is a genuine competitor. The path-aware policy finished with the lowest aggregate mapping error (0.1778) at mid-pack energy, calibration coverage of 0.942, and zero false confirmations.
Aggregate mapping error against mean energy for all six policies, one point each, from the single blind run. The path-aware policy reaches the lowest error at mid-pack energy. Hover a point for its numbers.
What was learned
Three findings now shape the platform. First, when sensors measure continuously while moving, the path itself is a measurement, so acquisition must value information along the entire route as well as at the destination. Second, results tuned on few random worlds rarely survive many; every calibration now passes a wider gate before it is trusted. Third, the remaining detection gap is a coverage problem: it points the next round of work at where the vehicle goes rather than at how it reasons.
Limits and what's next
The result has clear edges. The policy won the aggregate but fell one world short of per-world dominance, winning 14 of 20 against the strongest baseline where the registered bar was 15. A specialist anomaly hunter still detects more at this budget (0.800 vs 0.767), a coverage gap that is the next lever under test. And two improvements we believed in, deeper lookahead and objective-conditioned acquisition weights, were refuted by their own evaluation gates and ship disabled.
The next stage is physical: an instrumented autonomous surface vehicle running the same controller on the real lake, behind state research permits and the same deterministic safety gateway. Success criteria will be registered before the first mission, and the results will be published here, wins and losses alike.