One SoC, Multiple Engines, Matching Results
Taking a multicore workload through Dancer and checking its result and waveform streams across execution environments.
A multicore workload becomes much more interesting when it has to survive a change in execution engine.
The processors must still execute the same firmware. DMA must copy the same bytes. Interrupts must arrive and complete with the same meaning. Moving the workload must not quietly replace the behavior being checked.
We took the four-core development SoC from the first post in this series through Dancer, the execution infrastructure we are building for agent-operated hardware workflows. The recorded run reproduced the hosted result and waveform streams byte-for-byte on a bare-metal execution host.
The subject remained the same interacting SoC model: four cores, shared memory, multiple clock domains, four DMA channels, timers, serial output, GPIO, and external interrupts.
Keep the Workload Intact
The development scenario already contained meaningful work: firmware-driven peripheral transactions, deliberately unaligned DMA copies, intermediate checks, interrupt handling, and independent destination-memory readback.
Dancer consumed that same shared scenario and its assertions. A shorter constant-memory workload was not substituted for the multicore behavior.
That continuity matters. If the stimulus, initialization, or expected behavior changes while moving between tools, matching results can become difficult to interpret. The question we wanted to answer was whether the execution paths agreed on the same experiment.
- 01Defined SoC workload
Firmware traffic, initialization, intermediate checks, and expected outcomes.
- 02Checked execution
Reference evaluation and recording retain the relevant observations.
- 03Dancer replay
Hosted and bare-metal execution reproduce the selected result and waveform streams.
Make the Comparison About Observable Behavior
The workload began with controlled reset and memory initialization, then released the cores to execute. The checks followed behavior through requests, responses, memory effects, peripheral events, and final completion state.
The recorded four-core run included 1,699 accepted fabric requests and 77 peripheral transactions. DMA reached sixteen simultaneous end-to-end operations, and the scenario checked that outstanding work drained afterward.
Dancer's recording and replay path was checked against reference evaluation. The hosted and bare-metal runs were then compared using the complete selected result and waveform streams.
A successful run required more than a terminal completion message. The streams had to match, the run had to complete its recorded work, and runtime or transport errors could not be ignored.
Result + waveform streams
The hosted run provides the reference streams for the recorded workload.
Result + waveform streams
The validator checks exact equality and completion, including runtime and transport outcomes.
This gave us an inspectable chain from a defined workload to observations reproduced in another execution environment.
What the Recorded Run Exercised
The executable gate-level model contained 319,304 gates. The bare-metal run used a 32-thread host configuration and completed the recorded workload with no reported runtime or network error. Its result and waveform streams matched the hosted references exactly.
Those numbers describe different things:
| Quantity | What it describes |
|---|---|
| Four cores | Processors modeled inside the development SoC. |
| 319,304 model gates | The logical executable model used in this run, not placed-and-routed silicon area. |
| 32 host threads | The execution-host configuration, not the number of modeled processor cores or a measured speedup factor. |
| Byte-identical result and waveform streams | Agreement over the selected observations for this recorded experiment. |
Bare metal here describes the machine executing the simulator. The SoC itself remained a design model; this was not a test of a fabricated version of that chip.
Matching Engines and Checking Intent Are Different Jobs
Execution consistency is valuable, but two engines can reproduce the same design mistake.
The earlier processor example in the evidence post made that boundary concrete: matching outputs did not establish meaningful processor execution when the accepted stimulus had never enabled a pair to fetch an instruction.
For the multicore workload, firmware checks, intermediate assertions, interrupt-class counts, and independent memory readback supplied requirements beyond stream equality. Those checks explain why the observed activity mattered.
They do not amount to complete independent ISA validation, coherent-memory verification, exhaustive reset-under-load coverage, or physical timing signoff. The recorded checkpoint is an uncached development platform with one substantial concurrent workload.
This separation helps make the result useful: the workload checks address intended behavior, while cross-engine comparison examines whether execution preserves the observations being relied upon.
An Execution Engine for an AI-Native Workflow
Dancer's role in the broader Memdance stack is to make execution useful to investigation and iteration.
For an engineer or an agent, the loop should remain coherent: select the design and workload, run it, inspect the outcome, localize a disagreement, make a change, and repeat the relevant experiment. Results should remain connected to the conditions under which they were obtained.
- 01Run
Use an identified design and workload.
- 02Inspect
Understand the observations and any disagreement.
- 03Revise
Make a change tied to the engineering requirement.
- 04Replay
Recheck the result under known conditions.
The recorded run gives us a concrete foundation for that direction. The broader Dancer roadmap develops execution capacity, debugging, replay, and agent-operated workflows around consistent hardware behavior. Those are engineering goals being developed beyond this checkpoint, not capabilities established merely by one successful run.
As execution technology evolves, the product question remains practical: can engineers change how a workload runs while continuing to understand what ran, what was observed, and what the result supports?
From Generated Source to Repeatable Engineering
The series describes one connected effort:
- Build an interacting SoC with agent-operated engineering. Give the work a written architecture and executable requirements.
- Use evidence to guide corrections. Measure the behavior that matters and preserve the basis for each conclusion.
- Carry that workload through execution and replay. Check that its observations survive the handoff.
- Keep the experiment recoverable. Preserve useful work and observations across a planned interruption.
This is the direction of Memdance's AI-native hardware engineering stack. Generation, investigation, verification, and execution need to work together so that an agent's contribution can become a result an engineering team can understand and build on.
The goal is repeatable engineering progress: a design that runs, evidence that explains the result, and a workflow that can carry both forward.
In our next post, we’ll take the operational step: checkpoint a running experiment, power the host off, and resume with its results intact. Continue to Pause. Power Off. Resume the Same Experiment. →