← The journal

From Architecture to a Running Multicore SoC

An AI-operated workflow built on Memdance’s design, execution, and verification infrastructure.

Four processor cores were executing firmware. Four DMA channels were moving data. Timers, a serial transmitter, GPIO, and external interrupts were sharing the same system. Requests crossed clock boundaries, responses returned to their owners, and the workload eventually drained the outstanding work.

The milestone was not generated source. It was an executable system whose interacting behavior we could inspect, challenge, and verify.

The infrastructure is the product. The SoC is a demanding workload for it.

Starting from a written architecture and maintained building blocks, agents developed and revised the hardware source, operated our compilation and execution tools, investigated failures, and turned findings into corrections. Human direction established the architecture, goals, and acceptance criteria.

The design was compiled, executed, and checked through the technology we are building. The agent worked through that engineering environment throughout implementation and investigation.

Start with a System That Has to Interact

The development platform combined small RISC-V cores, shared uncached memory, a tagged request/response fabric, multiple clock domains, and a peripheral subsystem. The same parameterized source elaborated two-, four-, and eight-core configurations. The richer execution workload ran on four cores; the eight-core configuration exercised structural and capacity checks.

The integration had concrete obligations. Reset had to release in the right order. Firmware had to load and read back correctly. Responses had to reach the requesters that owned them. A stalled consumer could not turn a completed transaction into lost data. Interrupt handling had to finish the specific work that raised the interrupt.

These requirements made the design much more revealing than a collection of blocks that compiled separately.

An interacting development SoC
Four firmware-executing RISC-V cores
Memory

Shared 1 MiB memory

The workload loads, executes, and checks data through the shared uncached path.

Interconnect

Requests with identifiable owners

Traffic crosses clock boundaries and responses return to the correct requester.

Data movement

Four DMA channels

The recorded workload reaches sixteen simultaneous end-to-end operations.

Peripherals

Timers · UART · GPIO

Firmware handles the events while the scenario checks their distinct effects.

A functional map of the development design. It describes the modeled SoC, not the private implementation of the engineering tools.

The SoC integration was built from the architecture using maintained core and interface components. Reuse gave the team useful starting points; it did not remove the need to establish that the assembled system behaved correctly.

Put the Agent Inside the Engineering Loop

Our working model gave agents concrete feedback at each step: compile the design, exercise a focused behavior, inspect the result, identify the violated requirement, make a bounded change, and rerun the affected checks.

A compiler rejection could point to an invalid connection or inconsistent declaration. An execution result could show a missing response, an unexpected fault, or a workload that had not reached the required activity. Those are different engineering problems, and they call for different changes.

The architecture and expected outcomes supplied a stable target. The implementation could evolve without redefining success around whatever it happened to do.

Agent-operated engineering, checked by the infrastructure
  1. 01
    Define the requirement

    Architecture and expected behavior give the work a stable target.

  2. 02
    Develop the design

    Agents author and revise hardware and supporting workload source.

  3. 03
    Exercise and inspect

    Compilation and execution provide concrete feedback.

  4. 04
    Review and rerun

    Judge the change against the original requirement and affected checks.

The engineering result comes from the complete loop. Source generation is one activity within it.

The agent needs concrete engineering facts, executable checks, and bounded changes tied to the failure being addressed. Human review can then focus on consequential decisions and acceptance, with the investigation available for inspection.

Firmware Made the Integration Real

The four-core workload drove the peripherals through processor-authored memory-mapped traffic and firmware handlers.

Each core started a distinct, deliberately unaligned DMA copy. Completion handlers checked command identity and status, inspected copied data, completed the interrupt, and returned. The scenario later halted instruction fetch and independently read the destination memory, including the edge words affected by partial byte writes.

Core zero also transmitted a serial byte, drove GPIO, handled a later GPIO edge, and serviced an external interrupt. Timers supplied work across the cores. The workload checked interrupt classes separately so a missing GPIO event could not be disguised by an extra timer event.

The recorded run included:

Observation Recorded result Why it matters
Executing processor cores Four Work came through firmware and the integrated system.
Fabric requests accepted 1,699 The run exercised a substantial sequence of routed interactions.
Peripheral transactions 77 Peripheral access was part of the running workload.
Peak end-to-end DMA operations 16 All four channels reached four simultaneous operations.
Completion state Outstanding work drained; checked fault outputs clear Reaching a peak was followed by checked completion.

These are observations from one defined workload, not throughput measurements. Their value is that they say what the system actually exercised.

The Blocks Existed. The Behavior Still Had to Be Earned.

One early integration problem was easy to miss on a block diagram. Four DMA engines each had four tags, but a shared router kept only one request active until its response. More tags did not make that path concurrent.

The workload exposed the mismatch between declared capacity and reachable behavior. The routing path was corrected, and the resulting run reached the required sixteen-operation high-water mark before draining the work.

Another problem initially looked like an interrupt-path defect. A GPIO edge was applied before firmware had armed the input, despite the timestamp making it appear late enough. An explicit arming checkpoint separated a scenario-timing mistake from a hardware bug.

Both findings depended on asking more precise questions than whether the program had reached a terminal loop. The next post examines how evidence and measurement guided those decisions.

Carry the Same Workload into Another Execution Engine

Once the behavior was established in the development execution path, we carried the same scenario and its intermediate checks into Dancer, our simulation and emulation technology.

The handoff retained the actual multicore workload. It did not replace DMA, peripherals, and interrupt traffic with a smaller demonstration simply to obtain matching output.

That provided another useful question: would a different execution path reproduce the observations that mattered? The third post follows that work through Dancer and replay.

What This Milestone Says About the Infrastructure

This checkpoint produced an interacting, uncached development SoC model. It did not close coherent-cache verification, broad independent architectural conformance, sustained acceptance-scale traffic, or physical chip implementation.

The progress is still concrete: parameterized source, executing firmware, concurrent traffic, independently checked memory results, and a workload that could move between execution paths while keeping its checks.

The agent-operated workflow carried the work through implementation, investigation, and validation. The executable requirements supplied a stable basis for deciding whether a change was progress.

Memdance is building the infrastructure for that way of working. Our ambition is to let intelligent agents participate in larger portions of chip development while giving engineers a clearer basis for understanding, reviewing, and directing the result.

The milestone was a running system that could explain its behavior—not merely a generated design.

In our next post, we’ll examine how built-in evidence and measurement exposed problems a successful-looking run could miss—and helped the agent-operated workflow choose the right correction. Continue to Evidence Built Into the Engineering Loop →