← The journal

Evidence Built Into the Engineering Loop

How observed concurrency, event history, and meaningful execution checks turn results into better engineering decisions.

Four DMA engines. Four tags per engine. Sixteen operations available on paper.

Yet the shared routing path allowed only one request to remain active until its response arrived.

The declarations looked capable. The observed behavior exposed the limitation.

Memdance is building the engineering infrastructure that makes such distinctions available to agents. The SoC is a workload for that infrastructure. Evidence becomes useful when it changes the next engineering decision: where to investigate, what to correct, which result still applies, and what needs to be checked again.

Measure the Behavior the Requirement Actually Describes

The multicore SoC in the first post needed to demonstrate sixteen simultaneous end-to-end DMA operations. Counting allocated tags or inspecting queue declarations could not establish that result.

The workload therefore observed both per-channel and global high-water marks. It checked requests and responses, completion state, destination data, and the absence of leftover ownership or fault state.

That exposed a real integration limit. The shared router serialized the path despite the surrounding outstanding-transaction capacity. Correcting the routing behavior allowed the workload to reach four operations in every channel and a global peak of sixteen, then drain the system.

Configured capacity must become reachable behavior
Declared capacity

4 channels × 4 tags

Sixteen available identities do not establish sixteen simultaneous operations through the complete path.

Integration bottleneck

One active routed request

The shared path held a request until its response, limiting the concurrency the workload could exercise.

After correction: peak 16 end-to-end operations · checked completion and drain
The pre-correction observation concerns the shared routing path. The peak of sixteen is measured in the corrected workload.

The important measurement was not a larger number by itself. It was a number tied to a requirement, reached by executable traffic, and followed by checked completion.

This is the kind of relationship we want engineers and AI agents to be able to inspect directly.

A Missing Interrupt Can Be a Scenario Bug

The same SoC supplied another useful example. A GPIO input change appeared to be scheduled after the point at which firmware should have armed the input. In fact, the scenario applied that input over the interval leading to the timestamp. The edge arrived too early.

Timer and DMA traffic made a synchronizer or interrupt gateway look like a plausible suspect. Inspecting the conditions around the event led to a different correction: establish an explicit low-input arming checkpoint, check the firmware-visible GPIO configuration there, and apply the rising edge in the next interval.

The workload also counted interrupt classes separately:

Interrupt source Expected claims in the workload
External input 1
DMA 4
UART 1
GPIO 1
Timer 4

An aggregate count of eleven could hide the wrong combination. Checking the classes made the claim more precise.

The evidence helped distinguish an error in the test's timing from an error in the design. That is valuable whether the next investigator is a person or an agent.

An Earlier Processor Experiment Made the Lesson Sharper

In a separate, earlier processor-cluster validation effort, replay outputs matched—but the accepted scenarios had never enabled a processor pair to fetch an instruction.

The comparison established agreement on the recorded activity. It did not establish that a program could execute.

When a faithful executing workload was introduced, both cores in a released pair took the same illegal-instruction trap on their first fetch. The investigation found that an instruction response was selecting the wrong memory data. Additional checks exposed maintenance-response and error-correction handling defects.

A controlled probe helped separate the faulty execution path from the lockstep comparison machinery. Corrections were then checked with actual instruction retirement, maintenance readback, and fault scenarios. The subsequent validation sequence passed across the recorded execution environments.

An earlier processor-cluster investigation
  1. 01
    Agreement on a limited run

    The accepted stimulus had not enabled the pair to fetch an instruction.

  2. 02
    Exercise actual execution

    The first fetch exposed a common failure in both cores.

  3. 03
    Correct and verify

    Retirement, readback, and fault checks exercised the corrected behavior.

A separate earlier development experiment. Replay agreement and meaningful processor execution answer different questions.

The earlier green result was retained with its narrower meaning. The correction required new evidence that exercised the behavior previously missing from the test.

This was a development investigation, not a finding from fabricated silicon. Its lesson is broadly useful: agreement becomes meaningful only when the experiment has exercised the behavior behind the claim.

Evidence Should Help Decide What to Do Next

An AI engineering workflow needs more than a collection of reports. It needs results that remain connected to the design, conditions, and requirements they concern.

That connection lets an investigator ask useful questions:

  • Did this run actually reach the concurrency the design promises?
  • Did the failure originate in hardware behavior, a test assumption, or the execution environment?
  • Does this passing result apply to the design after the proposed change?
  • Which relevant behavior remains unchecked?

We are designing Memdance so these relationships are part of the engineering workflow. An agent should be able to follow a failure to its context, propose a reviewable correction, and judge the rerun against the same intended behavior.

The value extends to collaboration. Another engineer or agent should be able to understand why a decision was made without inheriting an entire conversation or trusting an unexplained green check.

Measurement Belongs Beside Correctness

Performance and resource measurements need the same discipline. A shorter workload can make a run faster without improving the execution engine. An optimization can reduce one cost while changing behavior or moving work elsewhere.

A useful comparison records what was exercised, what was measured, and which conditions were held constant. Correctness checks establish whether the proposed improvement preserved the required behavior; measurements establish the observed tradeoff.

Our goal is for agents to use both when evaluating alternatives. They should be able to explain an engineering choice in terms of behavior and measured consequences, rather than the plausibility of the proposed source change.

Two kinds of evidence behind an engineering choice
Correctness

Did the required behavior survive?

Use explicit expectations, relevant corner cases, and the scope of the checks.

Measurement

What changed in the cost of doing the work?

Keep workload, operating conditions, and the quantity being measured visible.

Investigate · Compare · Correct · Recheck
The product direction is to make both kinds of information useful to engineers and agents throughout the workflow.

Build Confidence That Survives the Next Change

A result remains valuable when someone can inspect its basis, reproduce the relevant behavior, and recognize when it no longer supports a new conclusion.

That is why our work treats design meaning, execution results, measurements, and verification evidence as connected concerns. The examples above demonstrate concrete benefits in bounded development work; the broader product ambition is to make that discipline natural throughout the engineering loop.

Evidence does not make every conclusion correct automatically. It gives the workflow a better basis for finding a wrong conclusion and correcting it without losing the history of what happened.

For customers, that means a clearer route from a symptom to an engineering decision. For engineers, it means more useful investigations and reviews. For AI, it provides the context needed to take actions that can be checked and challenged.

In our next post, we’ll introduce our simulation and emulation system, Dancer, and follow the same SoC workload through checked execution and replay. Continue to One SoC, Multiple Engines, Matching Results →