Tracing a Hardware Failure to Evidence
An independent model, a controlled correction, and a result with clear boundaries.
What does an AI engineering agent need in order to investigate a hardware failure reliably?
Modern hardware development already has sophisticated tools for simulation, formal verification, debugging, waveform analysis, synthesis, and design navigation. At Memdance, we are exploring a narrower question: What changes if an AI agent can work with structured semantic and verification information maintained by the engineering system, instead of repeatedly reconstructing design meaning from source code, logs, waveforms, and tool output?
This experiment exercises one part of a broader system. Memdance is building an integrated AI-native stack spanning design representation, compilation and transformation, execution, netlist generation, verification, evidence, and agent-driven engineering workflows. The FIFO case study focuses specifically on whether design meaning and verification evidence remain connected as hardware moves through that stack.
The FIFO is a test vehicle for the infrastructure. The result we are evaluating is an automated investigation that connects observed behavior, the relevant requirement, and a checked correction.
To test that idea, we built a deliberately small experiment around a FIFO. The FIFO itself is not particularly interesting—the experiment is. We wanted to determine whether an observed implementation-level failure could be connected back to the correct design behavior and associated verification requirement, and whether that connection would survive transformations and deliberate attempts to confuse the localization process. We also wanted verification evidence to retain its precise scope rather than silently becoming a stronger correctness claim than the evidence justified.
The Setup & Mismatch Detection
The test contains a small synchronous FIFO with an intentional defect in its read/write behavior. Alongside the implementation, we run an independent queue model that knows only the externally visible FIFO protocol.
The two-entry FIFO held [1, 0]. A simultaneous dequeue and enqueue of 0 reused the same memory address. The required result was the oldest queued value, 1; the defective collision policy returned the newly written 0 instead. The independent queue model found this shortest counterexample by enumeration.
Lowered FIFO netlist
Independent queue model
That separation is important. Two execution engines can agree perfectly with each other while faithfully executing the same incorrect design; execution agreement therefore does not establish that the implementation satisfies its intended behavior. The independent queue model provides a separate reference to answer: Given this sequence of FIFO operations, what should the externally visible result have been?
During the experiment, the implementation produced result[0] = 0 while the independent queue model expected result[0] = 1 at time t=3. At this point, we know only that the observed implementation behavior disagrees with the independent specification. We do not yet know which part of the design produced the behavior, which requirement is relevant, whether the relationship survives transformations, or whether a proposed fix actually addresses the root defect.
From Behavior to Design Contract
The system first localizes the failing implementation behavior back to the relevant semantic region of the design. Rather than reconstructing relationships solely through textual similarity, the identity chain is maintained directly by the compiler. An autonomous engineering agent should ideally be able to answer "What design behavior produced this observation?" without depending on convenient signal names, source locations, or object ordering.
- 01Observed failure
- 02Semantic localization
- 03Design behavior
- 04Verification requirement
Finding the part of the design responsible for an incorrect value explains what produced the behavior, but it does not explain why that behavior is incorrect. The experiment therefore connects the localized behavior to its associated verification requirement. For autonomous engineering, this distinction is critical: a useful investigation must connect an observation not only to implementation structure, but also to the engineering intent that gives that observation meaning.
Stress-Testing the Localization Engine
A localization mechanism can easily appear semantic while actually relying on structural coincidence—such as picking the first plausible memory, relying on fixed signal names, or matching source file line numbers. To challenge this, we deliberately introduced adversarial conditions into the design environment:
- Decoy Insertion: Placed a structurally similar decoy FIFO earlier in the hierarchy.
- Hierarchy & Naming Shifts: Added intermediate hierarchy levels, randomized internal net identifiers, and shifted source file offsets.
- IR Transformation: Regenerated lower-level object ordering and net mappings.
| Perturbation | Result |
|---|---|
| Internal names changed | Correct target retained |
| Source positions shifted | Correct target retained |
| Extra hierarchy introduced | Correct target retained |
| Similar decoy inserted first | Decoy rejected |
| Internal ordering changed | Correct target retained |
The target moved away from its original hierarchy and ordering while a structurally similar decoy occupied the earlier position. Localization still converged on the intended design behavior.
Both FIFOs received the same stimuli. The correction left the unrelated FIFO’s source unchanged, and the affected requirement was located again after recompilation so acceptance used fresh evidence for the repaired design.
Controlled Repair & Evidence Scoping
After retrieving the verification requirement associated with the affected behavior, the experiment evaluated a deliberately constrained set of three permitted candidate corrections. The candidates changed the read-during-write behavior: return the newly written value, hold the previous read output, or return the pre-write memory contents. This demonstration is not an unconstrained LLM inventing arbitrary hardware changes; it tests whether a reproducible failure can be turned into a structured investigation whose candidate corrections are judged by independent evidence.
The loop was deterministic: no LLM took part in detection, localization, candidate selection, or replay. The FIFO contract and independent queue oracle were supplied explicitly.
| Candidate | Original failure | Broader checking | Decision |
|---|---|---|---|
| A · Newly written value | Fail | Fail | — |
| B · Hold previous output | Fail | Fail | — |
| C · Pre-write memory contents | Pass | Pass | Accepted |
Replay & Bounded Exploration (4,096 Traces)
After applying Candidate C, the original counterexample passed in both netlist and direct execution. The system then explored the complete finite input space of all 4,096 four-event queue traces.
- Before Repair: Counterexample found, retained, and replay-verifiable.
- After Repair: Original failure passed, and all 4,096 traces satisfied the independent queue model.
We explicitly do not claim the FIFO is universally proven correct; we claim that no counterexample was found in the complete finite search space defined for this experiment. The requirement remains marked as open, with the new evidence explicitly recorded as bounded. A bounded result should remain bounded, and a proof should be called a proof only when the relevant obligation has actually been discharged.
Evidence Integrity & Reproducibility
Automated systems can produce a particularly dangerous failure mode: producing a confident conclusion supported by stale evidence that no longer applies to the modified design.
- Stale Evidence Rejection: We deliberately attempted six forms of stale or tampered artifact reuse (presenting pre-repair state or mismatched validation info). All six were rejected.
- Deterministic Investigation: The baseline experiment was repeated independently, and all 30 retained artifacts reproduced byte-for-byte.
Reproducibility
30 artifacts
Reproduced byte-for-byte
Evidence integrity
6 stale or tampered cases
All rejected
Summary of Findings & Boundaries
To maintain engineering rigor, we explicitly separate what this experiment demonstrates from what remains unaddressed:
What This Experiment Demonstrates
- Semantic Localization: Implementation-level failures connect reliably to design behavior without relying on string matching.
- Adversarial Resilience: Identity chains survive signal renaming, hierarchy shifts, line offsets, and structural decoy insertion.
- Contract Binding: Observed defects connect directly to actionable engineering assertions.
- Independent Acceptance & Replayability: Corrections are judged by external models, and counterexamples remain part of the retained evidence suite.
- Scoped Evidence & Determinism: The six tested stale or tampered artifacts were rejected, bounded results stayed bounded, and 30 artifacts reproduced byte-for-byte.
What This Experiment Does Not Demonstrate
The demonstrated transformations were hierarchy flattening and netlist lowering. The experiment does not establish tracing through arbitrary retiming, technology mapping, or optimization that deletes structure. The fixture also includes reference state for its ordering assertion; it is not a claim about a minimal FIFO implementation.
- Scalability to 10-Billion-Transistor SoCs: Demonstrates identity preservation across IR transformations, not large-scale SoC footprint management or multi-gigabyte netlist traversal.
- Engineering Overhead: Maintaining structured semantic and verification information has computational and storage costs. One question we are evaluating is whether that additional machinery materially reduces the cost of automated investigation and verification.
- Replacement of Commercial EDA Toolchains: Memdance complements established signoff engines from Synopsys, Cadence, or Siemens by exporting standard artifacts (SystemVerilog) rather than attempting a total replacement.
- Arbitrary Repair Synthesis & Physical Signoff: Does not cover unconstrained AI synthesis, layout placement, timing closure, PPA optimization, or general specification completeness.
Why This Matters for AI Agents
Experienced hardware engineers already have sophisticated ways to trace drivers, navigate hierarchy, inspect waveforms, and perform formal analysis. However, the economics change when the primary consumer of that context is an autonomous agent.
Text-centric agent workflow
- 01Unstructured logs / waveforms
- 02Probabilistic text parsing
- 03Fragile assumptions
Structured engineering workflow
- 01Structured engineering state
- 02Direct programmatic query
- 03Deterministic tracing
For an agent, there is a fundamental difference between parsing raw text logs and asking structured programmatic queries: What design behavior produced this observation? Which engineering requirement governs it? What evidence supports it, and does that evidence belong to the current design state?
We believe maintaining a dependable connection between observed hardware behavior, design meaning, and verification evidence can provide a stronger substrate for increasingly autonomous hardware engineering.