Evidence Built Into the Engineering Loop
How observed concurrency, event history, and meaningful execution checks turn results into better engineering decisions.
Four DMA engines. Four tags per engine. Sixteen operations available on paper.
Yet the shared routing path allowed only one request to remain active until its response arrived.
The declarations looked capable. The observed behavior exposed the limitation.
Memdance is building the engineering infrastructure that makes such distinctions available to agents. The SoC is a workload for that infrastructure. Evidence becomes useful when it changes the next engineering decision: where to investigate, what to correct, which result still applies, and what needs to be checked again.
Measure the Behavior the Requirement Actually Describes
The multicore SoC in the first post needed to demonstrate sixteen simultaneous end-to-end DMA operations. Counting allocated tags or inspecting queue declarations could not establish that result.
The workload therefore observed both per-channel and global high-water marks. It checked requests and responses, completion state, destination data, and the absence of leftover ownership or fault state.
That exposed a real integration limit. The shared router serialized the path despite the surrounding outstanding-transaction capacity. Correcting the routing behavior allowed the workload to reach four operations in every channel and a global peak of sixteen, then drain the system.
4 channels × 4 tags
Sixteen available identities do not establish sixteen simultaneous operations through the complete path.
One active routed request
The shared path held a request until its response, limiting the concurrency the workload could exercise.
The important measurement was not a larger number by itself. It was a number tied to a requirement, reached by executable traffic, and followed by checked completion.
This is the kind of relationship we want engineers and AI agents to be able to inspect directly.
A Missing Interrupt Can Be a Scenario Bug
The same SoC supplied another useful example. A GPIO input change appeared to be scheduled after the point at which firmware should have armed the input. In fact, the scenario applied that input over the interval leading to the timestamp. The edge arrived too early.
Timer and DMA traffic made a synchronizer or interrupt gateway look like a plausible suspect. Inspecting the conditions around the event led to a different correction: establish an explicit low-input arming checkpoint, check the firmware-visible GPIO configuration there, and apply the rising edge in the next interval.
The workload also counted interrupt classes separately:
| Interrupt source | Expected claims in the workload |
|---|---|
| External input | 1 |
| DMA | 4 |
| UART | 1 |
| GPIO | 1 |
| Timer | 4 |
An aggregate count of eleven could hide the wrong combination. Checking the classes made the claim more precise.
The evidence helped distinguish an error in the test's timing from an error in the design. That is valuable whether the next investigator is a person or an agent.
An Earlier Processor Experiment Made the Lesson Sharper
In a separate, earlier processor-cluster validation effort, replay outputs matched—but the accepted scenarios had never enabled a processor pair to fetch an instruction.
The comparison established agreement on the recorded activity. It did not establish that a program could execute.
When a faithful executing workload was introduced, both cores in a released pair took the same illegal-instruction trap on their first fetch. The investigation found that an instruction response was selecting the wrong memory data. Additional checks exposed maintenance-response and error-correction handling defects.
A controlled probe helped separate the faulty execution path from the lockstep comparison machinery. Corrections were then checked with actual instruction retirement, maintenance readback, and fault scenarios. The subsequent validation sequence passed across the recorded execution environments.
- 01Agreement on a limited run
The accepted stimulus had not enabled the pair to fetch an instruction.
- 02Exercise actual execution
The first fetch exposed a common failure in both cores.
- 03Correct and verify
Retirement, readback, and fault checks exercised the corrected behavior.
The earlier green result was retained with its narrower meaning. The correction required new evidence that exercised the behavior previously missing from the test.
This was a development investigation, not a finding from fabricated silicon. Its lesson is broadly useful: agreement becomes meaningful only when the experiment has exercised the behavior behind the claim.
Evidence Should Help Decide What to Do Next
An AI engineering workflow needs more than a collection of reports. It needs results that remain connected to the design, conditions, and requirements they concern.
That connection lets an investigator ask useful questions:
- Did this run actually reach the concurrency the design promises?
- Did the failure originate in hardware behavior, a test assumption, or the execution environment?
- Does this passing result apply to the design after the proposed change?
- Which relevant behavior remains unchecked?
We are designing Memdance so these relationships are part of the engineering workflow. An agent should be able to follow a failure to its context, propose a reviewable correction, and judge the rerun against the same intended behavior.
The value extends to collaboration. Another engineer or agent should be able to understand why a decision was made without inheriting an entire conversation or trusting an unexplained green check.
Measurement Belongs Beside Correctness
Performance and resource measurements need the same discipline. A shorter workload can make a run faster without improving the execution engine. An optimization can reduce one cost while changing behavior or moving work elsewhere.
A useful comparison records what was exercised, what was measured, and which conditions were held constant. Correctness checks establish whether the proposed improvement preserved the required behavior; measurements establish the observed tradeoff.
Our goal is for agents to use both when evaluating alternatives. They should be able to explain an engineering choice in terms of behavior and measured consequences, rather than the plausibility of the proposed source change.
Did the required behavior survive?
Use explicit expectations, relevant corner cases, and the scope of the checks.
What changed in the cost of doing the work?
Keep workload, operating conditions, and the quantity being measured visible.
Build Confidence That Survives the Next Change
A result remains valuable when someone can inspect its basis, reproduce the relevant behavior, and recognize when it no longer supports a new conclusion.
That is why our work treats design meaning, execution results, measurements, and verification evidence as connected concerns. The examples above demonstrate concrete benefits in bounded development work; the broader product ambition is to make that discipline natural throughout the engineering loop.
Evidence does not make every conclusion correct automatically. It gives the workflow a better basis for finding a wrong conclusion and correcting it without losing the history of what happened.
For customers, that means a clearer route from a symptom to an engineering decision. For engineers, it means more useful investigations and reviews. For AI, it provides the context needed to take actions that can be checked and challenged.
In our next post, we’ll introduce our simulation and emulation system, Dancer, and follow the same SoC workload through checked execution and replay. Continue to One SoC, Multiple Engines, Matching Results →