← The journal

Pause. Power Off. Resume the Same Experiment.

A checkpointed Dancer run on Orchestra: resume execution after power-off and verify the completed result.

Partway through a 39-workload execution run, we requested a checkpoint and powered the machine off.

After a cold boot, the runtime resumed from the saved position. The capture receiver continued from the data it already held. At completion, the result and waveform streams matched the hosted references byte-for-byte.

The machine had stopped. The experiment did not have to start over.

The first three posts followed a multicore design through implementation, evidence-guided investigation, and Dancer execution. This fourth part turns to the operational infrastructure around that work: keeping an experiment recoverable, observable, and useful across interruption.

Running the Experiment Is Part of the Engineering

An agent that can investigate a design still needs to operate the work. It has to start the right experiment, observe progress, distinguish completion from interruption, and retrieve a result that belongs to the run it intended to perform.

For a long-running job, interruption can create several different problems. Execution may stop. A result transfer may disconnect. A machine may need to shut down. Restarting the command does not, by itself, establish which work survived or whether the resulting record is complete.

We are building infrastructure that makes those distinctions available to the workflow. The hardware workloads exercise that infrastructure; the ability to operate and recover the engineering work is part of the product.

The Orchestra Run

We exercised this lifecycle on Orchestra, the bare-metal execution environment used for these runs. Dancer's execution runtime and control tooling supplied the checkpoint, capture, and resume behavior.

This was a separate operational experiment using a 39-workload processor bundle, rather than another run of the four-core SoC from the earlier posts. It tested a complementary property: whether an interrupted experiment could continue and still produce the expected complete result.

The interruption was deliberate. We requested a pause, waited for the durable-checkpoint confirmation, and then powered the host off. After reboot, the runtime reported that it had resumed from the recorded position. The capture continued forward rather than starting a new history.

A controlled lifecycle interruption
  1. 01

    Run

    Begin the defined 39-workload experiment.

  2. 02

    Checkpoint

    Request a pause and confirm that progress has been saved.

  3. 03

    Power off

    Shut the host down after the checkpoint is durable.

  4. 04

    Cold boot and resume

    Continue from the saved execution position.

  5. 05

    Complete and compare

    Match the final result and waveform streams to the hosted references.

This was a planned, checkpointed shutdown on one host. It was not an abrupt power-loss or live-migration test.

That sequence is more informative than simply launching the same job twice. It exercised the boundary between running work, saved progress, a powered-off machine, and resumed execution.

Resume the Results, Too

Saving execution state is only part of the problem. An engineer also needs the observations produced before and after the interruption.

The receiver attached while the experiment was already running. After the power cycle, it resumed from its retained partial capture without downloading the entire record again. The completed capture contained approximately 475 MB of recorded data.

A separate retrieval of a selected portion matched the independently received capture. That provided an additional check on the path used to recover results.

Execution and observations must both survive
Runtime

Resume the work

Continue from the saved execution boundary rather than beginning a new run.

Capture receiver

Resume the record

Continue from retained partial data rather than retrieving the entire capture again.

Completed capture: approximately 475 MB · checked result continuity
The complete result and waveform streams matched the hosted references; a separately retrieved portion matched the received capture sample.

The useful outcome was continuity of both the work and its record. An agent could return to a completed experiment with its observations intact, rather than inherit an unexplained gap or silently mix two attempts.

Check Completion Against the Result

A successful restart message was not the acceptance condition. The run had to complete, and its final result and waveform streams had to match the hosted references generated from the same workload image.

Check Recorded outcome What it established
Checkpoint before shutdown Durable checkpoint confirmed The planned interruption followed a saved execution boundary.
Cold-boot resume Continued from the saved position The runtime resumed the interrupted run.
Capture reconnection Continued from retained partial data The receiver preserved continuity across the interruption.
Final result and waveform streams Byte-identical to hosted references The selected observable result was preserved through the lifecycle.
Separate result retrieval Matched the received capture sample The recovered data agreed across the checked retrieval paths.

These checks answer different questions. Reaching a final status does not establish stream equality. Matching a retrieved sample does not establish that execution completed. The combination gives the workflow a much more useful account of what happened.

Why This Matters for AI-Operated Engineering

When an agent owns implementation and investigation work, operational uncertainty becomes engineering uncertainty. It needs to know whether it should continue waiting, resume an interrupted experiment, investigate a failure, or evaluate a completed result.

A recoverable execution workflow gives it a firmer basis for those decisions. Engineers can also inspect the same record when reviewing a change or accepting a result.

The client-facing value is practical: preserve useful progress, recover the observations needed for diagnosis, and keep the connection between a run and its conclusions intact. The recorded experiment makes those properties concrete for this workload and host.

Dancer provides the execution behavior in this example. The Orchestra environment lets us exercise the operational lifecycle around it. Together, the result advances the broader Memdance goal: infrastructure through which agents can carry out engineering work that remains inspectable and repeatable.

An operational result, not just a restart message
Engineering workflow

Know what happened

Distinguish running, interrupted, resumed, and completed work.

Result review

Keep the conclusion inspectable

Connect the completed result to the workload and retained observations.

The product value is a workflow an agent can operate and an engineering team can examine.

What the Interruption Tested

This was a controlled, checkpointed shutdown followed by a cold-boot resume on the same host. It was not an abrupt power-loss test, a machine-to-machine migration, or a demonstration that every possible runtime failure is recoverable.

The preserved equality concerns this recorded workload's result and waveform streams. It does not independently prove the hardware design correct, establish physical chip behavior, or turn the development environment into a production service.

Those boundaries keep the result precise. The concrete accomplishment is that work and observations survived the planned lifecycle transition, and the completed output could still be checked against a known reference.

Four Parts of One Engineering Workflow

The series now connects four questions:

  1. Build: can agents use the infrastructure to implement and exercise an interacting design?
  2. Evidence: do the observations justify the engineering conclusion and next action?
  3. Execution: does the defined workload preserve its observed behavior across execution paths?
  4. Operations: can the work continue across interruption with its result still intact?

These are connected parts of the infrastructure Memdance is building. A design, a simulator, and a report become more useful when an agent can operate the complete workflow and an engineer can examine the basis for its decisions.

The experiment can stop without the engineering work losing its place.

Our earlier blinded debugging pilot examines a complementary question: can an agent investigate an unfamiliar failure, preserve correct hardware, and recognize when the evidence is insufficient? These system and operational milestones broaden the integration work; larger blinded evaluations remain ahead.