When the Right Fix Is No Fix
A blinded debugging pilot: repair a hidden defect, reject a bad expectation, and leave an unsupported physical claim unresolved.
A cancelled transaction changed a balance. The output looked fine. Four clock edges later, a probe finally exposed the corruption.
We gave a fresh, restricted AI solver the design, its behavioral contract, the input history, and access to Memdance's engineering tools. It did not receive the hidden mutation, intended repair, or answer key.
It found and repaired the defect. But that was only one of three problems in the pilot. In another, the hardware was correct and the test expectation was wrong. In the third, digital observations could not support the requested physical diagnosis.
One design needed a repair. One needed to be left alone. One question needed to remain unanswered.
The solver took those three different actions on its original attempts, with zero retries.
That distinction is the result we care about. An engineering agent needs to determine what the evidence warrants before deciding what to change.
What Was Actually Blinded?
Earlier experiments tested specific properties of our infrastructure with deliberately constructed failures. This pilot asked whether a solver without the author's diagnosis could use the available engineering context to investigate unfamiliar cases.
The protocol fixed the rubric, prompt, tools, budgets, compiler revision, and environment before the hidden cases were created. A separate author prepared the cases, and the complete eligible inventory was sealed before the solver began. All three generated cases were eligible; all three original attempts were retained.
| The solver could use | The solver could not see |
|---|---|
| Design source, behavioral contract, observed outputs, and complete digital input history | The mutation, intended repair, answer key, and healthy reference source |
| Restricted compiler queries, state observations, checkpoints, and replay | Private regression histories, their results, or the independent checker implementation |
| A bounded allowance of up to three guarded source edits per case | Author conversations, coordinator context, arbitrary files, shell access, or general network access |
No operator chose the solver's queries, supplied a diagnosis, or fed back hidden-test results during an attempt. The author, coordinator, tool gateway, execution environment, and model provider remained trusted. This separation tests investigation without an answer key; it is not a security claim about an untrusted operator or provider.
Three Problems, Three Different Engineering Decisions
| Problem | First-attempt outcome | Source edits | Independent final checks |
|---|---|---|---|
| Hidden cancellation defect | Located the faulty update, repaired it, and replayed the original history | 1 | 786 passed, 0 failed |
| Incorrect test expectation | Explained why the contract did not permit the expected behavior; preserved valid hardware | 0 | 803 passed, 0 failed |
| Unsupported physical claim | Reproduced the digital behavior; left the physical cause unresolved | 0 | 876 passed, 0 failed |
The check totals combine event-output checks over 20 fixed contract histories per case in two execution modes, hand calculations, and consistency checks. They are not counts of independent designs or discharged proof obligations. Passing the third case's digital checks does not establish the existence or cause of a physical pulse.
The two controls matter. A request to debug a failure can tempt an agent to assume that the hardware must be wrong. Here, one contract required reload behavior to take precedence over spending on the same edge. Changing the hardware to satisfy the supplied expectation would have broken that contract. In the physical case, explaining a digital latency did not justify inventing a pulse-level diagnosis.
The Defect Was Earlier Than the Symptom
The faulty design maintained two byte-sized balances, east and west, with one pending transaction. Cancellation should discard pending work without changing either committed balance. A probe copied a lane's pre-edge balance into a held report register.
West held 250. A transaction for 9 became pending. At cancellation, the faulty design added the amount anyway: (250 + 9) modulo 256 = 3. The report register concealed that change until a later west probe.
- 05
West transaction becomes pending
Committed balance: 250. Pending amount: 9.
- 06
Cancellation changes the balance
West incorrectly becomes 3. The held report does not reveal it yet.
- 07
Legitimate east-lane activity intervenes
Across edges 7–8, an east transaction opens and an east probe correctly reports 5.
- 10
The west probe exposes the defect
Expected: 250. Observed: 3. Four edges separate the mutation from its visible symptom.
- 11
The neighbor still behaves correctly
East again reports 5.
The solver replayed the complete history and observed west change at cancellation while east remained unchanged. Source inspection informed its hypothesis; compiler-provided context anchored the selected construct to the design being investigated. It then made one guarded correction to the west update condition, preserving the east condition.
After recompilation, the solver replayed the exact original input history. West stayed at 250 when the transaction was cancelled and reported 250 at the later probe. East continued to report 5. After the attempt ended, the independent validator checked the repaired artifact: the original 66 failed event-output checks became zero, and all 786 final checks passed.
Two Interventions That Must Not Be Confused
The solver also branched from a checkpoint just before cancellation. It replaced the cancel-only event with a commit-only event and kept the later history unchanged. Both executions produced a west balance and report of 3.
That established a useful, precise observation: on this history, cancellation behaved like commitment.
It did not establish the stronger counterfactual obtained by simply removing cancellation while leaving commit inactive. The author had performed that separate experiment, but the blinded solver did not. We do not credit the solver with it.
The successful source correction and original-history replay are additional evidence. They show that this targeted repair eliminated the observed corruption without damaging the checked neighbor. Keeping these results distinct makes the causal argument inspectable.
The Wrong Turns Stayed in the Record
The investigation was not a scripted sequence of successful tool calls. The solver submitted incomplete input frames, used an invalid query selector, and—in the physical control—requested a replay branch with an inconsistent history position. The tools rejected those requests. The solver corrected them without operator help.
Those failed requests remain in the retained attempts. Across the pilot, the solver made 28 tool calls and one source edit. Zero retries means that no completed attempt was discarded and replaced; it does not mean every request inside an attempt succeeded.
The recorded artifacts can be checked and the original and final digital tests rerun offline. That reproduction passed. It checks the retained evidence and behavior; it does not rerun the AI solver or establish that a new attempt would take the same path.
What This Pilot Establishes
A fresh restricted solver, without access to the mutation history or answer key, repaired a previously unseen design defect and handled two controls appropriately.
The investigation was source-assisted and compiler-anchored. Source text contributed materially to the reasoning; the solver did not perform a complete source-free traversal from symptom to cause. Finite histories supplied bounded regression evidence, not a universal proof. No physical silicon behavior was established.
This is a three-case feasibility result on small circuits. The controls make their contracts and evidence limits clear, so they do not establish resistance to sophisticated misleading evidence. The pilot supplies no general repair success rate or comparison with another toolchain.
The Infrastructure Behind the Decision
The ledger is a workload for Memdance's engineering infrastructure. The product ambition is the environment in which an agent can investigate a design, examine relevant state, make a bounded change, and check its consequences—and recognize when a change is unwarranted.
That connects this pilot to our work on evidence in the engineering loop. Useful automation needs more than a plausible patch. It needs a reason to believe the patch addresses the requirement, and a way to preserve uncertainty when the available observations cannot settle the question.
The next engineering question is how these capabilities compose in a larger, interacting system. Our four-part series follows a multicore SoC through implementation, evidence, execution, and operations. Broader blinded evaluations remain necessary to understand how reliably agents can handle unfamiliar problems at that scale.
Knowing what to change is part of engineering. Knowing what to preserve—and what remains unknown—is part of it too.