← The journal

One SoC, Multiple Engines, Matching Results

Taking a multicore workload through Dancer and checking its result and waveform streams across execution environments.

A multicore workload becomes much more interesting when it has to survive a change in execution engine.

The processors must still execute the same firmware. DMA must copy the same bytes. Interrupts must arrive and complete with the same meaning. Moving the workload must not quietly replace the behavior being checked.

We took the four-core development SoC from the first post in this series through Dancer, the execution infrastructure we are building for agent-operated hardware workflows. The recorded run reproduced the hosted result and waveform streams byte-for-byte on a bare-metal execution host.

The subject remained the same interacting SoC model: four cores, shared memory, multiple clock domains, four DMA channels, timers, serial output, GPIO, and external interrupts.

Keep the Workload Intact

The development scenario already contained meaningful work: firmware-driven peripheral transactions, deliberately unaligned DMA copies, intermediate checks, interrupt handling, and independent destination-memory readback.

Dancer consumed that same shared scenario and its assertions. A shorter constant-memory workload was not substituted for the multicore behavior.

That continuity matters. If the stimulus, initialization, or expected behavior changes while moving between tools, matching results can become difficult to interpret. The question we wanted to answer was whether the execution paths agreed on the same experiment.

Carry the same experiment across execution paths
  1. 01
    Defined SoC workload

    Firmware traffic, initialization, intermediate checks, and expected outcomes.

  2. 02
    Checked execution

    Reference evaluation and recording retain the relevant observations.

  3. 03
    Dancer replay

    Hosted and bare-metal execution reproduce the selected result and waveform streams.

The workload and its checks remain the basis for the comparison.

Make the Comparison About Observable Behavior

The workload began with controlled reset and memory initialization, then released the cores to execute. The checks followed behavior through requests, responses, memory effects, peripheral events, and final completion state.

The recorded four-core run included 1,699 accepted fabric requests and 77 peripheral transactions. DMA reached sixteen simultaneous end-to-end operations, and the scenario checked that outstanding work drained afterward.

Dancer's recording and replay path was checked against reference evaluation. The hosted and bare-metal runs were then compared using the complete selected result and waveform streams.

A successful run required more than a terminal completion message. The streams had to match, the run had to complete its recorded work, and runtime or transport errors could not be ignored.

Two environments, one recorded result
Hosted execution

Result + waveform streams

The hosted run provides the reference streams for the recorded workload.

Bare-metal execution host

Result + waveform streams

The validator checks exact equality and completion, including runtime and transport outcomes.

Recorded outcome: byte-identical streams
This establishes reproducibility for the selected observations. The modeled SoC was not a fabricated chip.

This gave us an inspectable chain from a defined workload to observations reproduced in another execution environment.

What the Recorded Run Exercised

The executable gate-level model contained 319,304 gates. The bare-metal run used a 32-thread host configuration and completed the recorded workload with no reported runtime or network error. Its result and waveform streams matched the hosted references exactly.

Those numbers describe different things:

Quantity What it describes
Four cores Processors modeled inside the development SoC.
319,304 model gates The logical executable model used in this run, not placed-and-routed silicon area.
32 host threads The execution-host configuration, not the number of modeled processor cores or a measured speedup factor.
Byte-identical result and waveform streams Agreement over the selected observations for this recorded experiment.

Bare metal here describes the machine executing the simulator. The SoC itself remained a design model; this was not a test of a fabricated version of that chip.

Matching Engines and Checking Intent Are Different Jobs

Execution consistency is valuable, but two engines can reproduce the same design mistake.

The earlier processor example in the evidence post made that boundary concrete: matching outputs did not establish meaningful processor execution when the accepted stimulus had never enabled a pair to fetch an instruction.

For the multicore workload, firmware checks, intermediate assertions, interrupt-class counts, and independent memory readback supplied requirements beyond stream equality. Those checks explain why the observed activity mattered.

They do not amount to complete independent ISA validation, coherent-memory verification, exhaustive reset-under-load coverage, or physical timing signoff. The recorded checkpoint is an uncached development platform with one substantial concurrent workload.

This separation helps make the result useful: the workload checks address intended behavior, while cross-engine comparison examines whether execution preserves the observations being relied upon.

An Execution Engine for an AI-Native Workflow

Dancer's role in the broader Memdance stack is to make execution useful to investigation and iteration.

For an engineer or an agent, the loop should remain coherent: select the design and workload, run it, inspect the outcome, localize a disagreement, make a change, and repeat the relevant experiment. Results should remain connected to the conditions under which they were obtained.

Execution as part of an engineering workflow
  1. 01
    Run

    Use an identified design and workload.

  2. 02
    Inspect

    Understand the observations and any disagreement.

  3. 03
    Revise

    Make a change tied to the engineering requirement.

  4. 04
    Replay

    Recheck the result under known conditions.

A consistent workflow is the product ambition across developing execution capabilities.

The recorded run gives us a concrete foundation for that direction. The broader Dancer roadmap develops execution capacity, debugging, replay, and agent-operated workflows around consistent hardware behavior. Those are engineering goals being developed beyond this checkpoint, not capabilities established merely by one successful run.

As execution technology evolves, the product question remains practical: can engineers change how a workload runs while continuing to understand what ran, what was observed, and what the result supports?

From Generated Source to Repeatable Engineering

The series describes one connected effort:

  1. Build an interacting SoC with agent-operated engineering. Give the work a written architecture and executable requirements.
  2. Use evidence to guide corrections. Measure the behavior that matters and preserve the basis for each conclusion.
  3. Carry that workload through execution and replay. Check that its observations survive the handoff.
  4. Keep the experiment recoverable. Preserve useful work and observations across a planned interruption.

This is the direction of Memdance's AI-native hardware engineering stack. Generation, investigation, verification, and execution need to work together so that an agent's contribution can become a result an engineering team can understand and build on.

The goal is repeatable engineering progress: a design that runs, evidence that explains the result, and a workflow that can carry both forward.

In our next post, we’ll take the operational step: checkpoint a running experiment, power the host off, and resume with its results intact. Continue to Pause. Power Off. Resume the Same Experiment. →