← The journal

Building a Chip Validation Stack from the Ground Up

An Arm validation stack for controlled execution, architectural checks, stress workloads, and precisely scoped results.

At Memdance, we are building tools and technology for both chip development and chip validation. Our work extends from creating and checking hardware designs to developing the systems needed to examine behavior on the eventual machine.

Apex is our work on the validation side: an Arm chip validation stack that brings together a bare-metal execution environment, architectural checks, stress workloads, and structured results. Its intended destination is the chip under test, where it can exercise behavior across cores, memory, interrupts, privilege levels, and architectural extensions.

The principle behind it is simple: a validator must know what it attempted, whether the prerequisites were satisfied, what it observed, and how strong a conclusion that observation supports. That principle guides both the tests and the environment that runs them.

Our development work has used QEMU. We have not yet booted Apex on physical bare-metal hardware. What follows describes the stack we have built, the validation areas we are developing, and what our development results establish.

Why control the execution environment?

A processor exposes much of its behavior through ordinary software. Some validation questions, however, require authority that an application does not have. Testing a protection boundary, changing an address mapping, or deliberately provoking an architectural fault requires control over the conditions surrounding the operation.

Apex owns enough of the execution environment to establish controlled conditions and observe architectural behavior directly. Its bare-metal design supports this work across privilege levels, cores, memory, and interrupts.

That control carries a responsibility. The validator must establish that its own environment is functioning before relying on it. A failure to prepare a test correctly cannot support the same conclusion as a correctly prepared test observing an invalid architectural result.

Portability also depends on discovering what a target actually provides. An architecture label alone does not describe every available extension, platform facility, or access restriction. Apex selects work according to the capabilities and prerequisites available to it.

The scope of Apex

The current catalog contains 103 workload descriptors at varying levels of maturity. This is an inventory of defined workloads, not a count of completed physical validation paths or validated features. Some paths execute architectural operations; others remain at the modeling or integration stage.

The work spans several broad areas:

Validation area Questions it addresses
Memory-system behavior Are shared data, atomic updates, ordering, and address translation behaving as required?
Privilege and protection Do permitted operations succeed and prohibited operations produce the expected response?
Architectural extensions Do specialized computations and their interactions with memory preserve the expected result?
Virtualization Are guest boundaries, translation contexts, and privileged behavior correctly enforced?
Interrupt delivery Are events delivered and handled with the required destination, ordering, and priority?
Reliability Can errors be observed, reported, and handled according to the relevant contract?
Performance monitoring Are the observations available and appropriate for understanding the exercised workload?

Apex spans Armv9 features and earlier AArch64 capabilities that remain fundamental to those systems. Selected headings use Arm architectural feature identifiers to connect the discussion to public terminology. These names identify the extension being discussed; they do not imply complete Apex coverage of that extension. See Arm's Feature names in A-profile architecture.

From execution to a meaningful result

For each test, we need a clear connection between its setup and its conclusion. The process has four parts: establish prerequisites, execute a controlled operation, observe the outcome, and interpret that outcome against an explicit expectation.

A test completing is only one observation. Depending on the scenario, we may need to know whether data remained consistent, whether an update became visible, whether an access was rejected, or whether an expected event arrived.

This is why our workloads are designed around explicit invariants and independently checkable outcomes. Activity becomes useful validation when there is a precise condition that can be violated and a way to recognize that violation.

It also gives a passing result a boundary. A pass means that the exercised case satisfied its checks under the recorded conditions. It does not establish every behavior of the feature or every possible interaction on the chip.

From execution to a meaningful result
  1. 01
    Prerequisites

    Capabilities and setup

  2. 02
    Execute

    Controlled operation

  3. 03
    Observe

    Record the outcome

  4. 04
    Interpret

    Compare with the expectation

A result is interpreted against the workload’s expectation and the conditions actually exercised.

Cache coherency and multicore behavior

Shared-memory software depends on cores agreeing about data as ownership moves between them. Apex includes workloads that exercise publication, consumption, reuse, and contention through familiar concurrent operations.

A shared work queue illustrates the expectation. Accepted work should remain intact and be consumed exactly once. A validation workload must be able to recognize missing work, duplicate consumption, and incorrect data while participants compete for shared resources.

The same principle applies to shared counters and object lifetimes. The useful question is whether the operation retains its meaning as multiple cores participate. A high level of traffic helps create challenging conditions, while the invariant supplies the correctness check.

Placement matters too. Interactions among nearby cores can exercise different parts of a system from interactions across a broader topology. Meaningful coverage must account for which participants actually took part.

Memory ordering

Memory ordering determines when one participant is entitled to observe another's work. A validator needs to distinguish an outcome allowed by the architecture from one that violates the synchronization used by the program.

Apex combines focused ordering checks with larger concurrent workloads. In a publication scenario, for example, the expectation depends on the synchronization establishing visibility. The test must interpret the observed data in that context.

Atomics — Large System Extensions (FEAT_LSE)

Atomic operations add a related question: do concurrent updates preserve the expected result when multiple participants contend for the same state? Apex includes workloads that exercise LSE atomics alongside other atomic mechanisms and check their interaction under load.

These checks connect architectural guarantees to the software patterns that depend on them, without assuming that every surprising interleaving is a hardware defect.

Address translation

Address translation sits between a program's view of memory and the physical storage it accesses. Validation needs to observe what an access actually reaches after a mapping changes.

Apex includes translation checks that examine whether memory accesses reflect the intended mapping after the required maintenance. Other scenarios address translation reuse and the architectural state associated with memory accesses.

The distinction is practical: software recording the desired mapping does not by itself establish that subsequent accesses use it. The observable result must agree with the completed transition. This becomes especially relevant when mappings change while other participants continue working.

Privilege and protection

Protection mechanisms need both positive and negative checks. A permitted operation should succeed. An operation that crosses the configured boundary should produce the expected architectural response.

Apex provides controlled execution for these scenarios, including checks involving memory protection and control flow. The outcome must be specific enough to distinguish an expected fault from an unrelated failure.

This makes setup part of the result. Before judging enforcement, the validator must establish that the relevant capability was available and the intended protection was active. Otherwise, a missing or unexpected fault has an ambiguous meaning.

The aim is to make each conclusion traceable to the boundary that was actually exercised, while keeping broader security claims separate from individual functional checks.

Memory tagging — MTE (FEAT_MTE)

Apex includes checks involving permitted tagged-memory operations and expected tag-mismatch faults. The validation question is whether the observed access behavior agrees with the protection that was established.

Pointer authentication — PAC (FEAT_PAuth)

Our pointer-authentication work examines valid authentication and the response to invalid pointers. Each result must be interpreted against the target's supported behavior and the conditions actually exercised.

Branch target identification — BTI (FEAT_BTI)

Apex includes control-flow checks involving permitted and prohibited branch destinations. The expectation depends on protection being active, and the observed exception must correspond to the intended check.

Architectural extensions and system behavior

Specialized computation and system facilities expand the kinds of behavior a validator must examine. Interrupt delivery and performance monitoring are part of this wider scope, alongside the named extensions below.

These areas are at different stages of development. Some remain at the modeling or integration stage, and platform support determines which scenarios can execute. The feature names describe the scope of the work, not a list of completed hardware-validation results.

Scalable vectors — SVE/SVE2 (FEAT_SVE, FEAT_SVE2)

Apex's vector work examines computation and data visibility across participating cores. Checks must account for the target's available vector capabilities and the synchronization under which results are observed.

Matrix compute — SME/SME2 (FEAT_SME, FEAT_SME2)

Our matrix-compute scenarios address numerical results and execution-state transitions. This work includes models and integration still requiring target execution, with physical validation ahead.

Reliability — RAS (FEAT_RAS)

Apex's reliability work addresses error observation and reporting. Meaningful injection and containment checks also require suitable hardware and firmware cooperation; those physical results remain ahead of us.

Resource partitioning — MPAM (FEAT_MPAM)

Resource-partitioning scenarios ask whether configured controls affect the intended participants. Establishing that result requires observations from a capable platform, beyond merely detecting the extension.

Enhanced Counter Virtualization — ECV (FEAT_ECV)

The catalog includes counter-isolation scenarios. Their conclusions depend on the virtualization environment actually exercised; a modeled result cannot establish enforcement on a physical target.

Fine-Grained Traps — FGT (FEAT_FGT)

Our trap-control scenarios concern whether selected accesses produce the expected response. These remain part of the developing virtualization coverage, with target access and setup as prerequisites.

Stress testing and burn-in

Stress testing is an explicit part of Apex. Repeated execution can sustain contention, revisit state transitions, and create more opportunities to observe a violation. The correctness checks remain active throughout the run.

We support profiles ranging from short development checks to sustained stress and extended burn-in workloads. Duration and load serve the test's purpose: repeatedly exercise behavior while preserving a meaningful expected outcome.

For example, an extended run can repeatedly transfer shared work, contend on atomic state, or change translations while checking the resulting accesses. Reports must distinguish completed work from an interrupted or inconclusive run.

We have no physical-hardware stress or burn-in results yet. Thermal characterization, power-state qualification, and aging conclusions require suitable platform controls and measurements in addition to repeated execution.

Structured results

A validation result should help an engineer decide what to investigate next. A single pass/fail indicator cannot express every reason a run might end.

Apex distinguishes an observed architectural violation from an invalid platform setup, a failure in the validator, an inconclusive timeout, or a test that could not run because a prerequisite was unavailable. Results retain the context needed to interpret the observation.

These distinctions prevent a failed launch or unavailable feature from being mistaken for evidence of defective silicon. They also keep an incomplete run from becoming an unjustified pass.

The reporting is intended to serve both human investigation and automated analysis. For automated consumers, the distinction between “failed,” “unsupported,” and “inconclusive” is as important as the observed value itself. In either case, the consumer needs to understand what ran, what was observed, and which conclusion is supported.

What QEMU has demonstrated

QEMU has allowed us to develop and exercise the stack before a successful physical boot. Our development runs have exercised multicore bring-up, selected memory-translation checks, controlled lower-privilege execution, expected return behavior, and bounded execution.

Those results establish progress in the validation software and its interaction with the emulated machine. They help us test whether the environment can prepare work, execute it, recover control, and report an outcome.

They do not establish complete coverage of every catalog entry. Emulated execution also has limits: a capability may be unavailable, modeled differently from the target conditions of interest, or insufficient for a particular observation.

A useful development result preserves those limits rather than treating all completed runs as equivalent.

The physical-hardware boundary

Apex has not yet run on physical bare-metal hardware. We therefore have no silicon-validation results, physical scaling measurements, or hardware burn-in results to report.

QEMU cannot provide evidence about a particular chip's physical cache and interconnect behavior, thermal conditions, or silicon defects. Those questions require execution and observation on the intended hardware under appropriate conditions.

The next hardware result must begin with a successful boot and a trustworthy test environment. Only then can individual observations support conclusions about the target.

Different environments, different evidence
Development evidence

QEMU

  • Multicore bring-up
  • Selected translation checks
  • Controlled execution and reporting
Work ahead

Physical hardware

  • Successful boot
  • Trustworthy test environment
  • Bounded hardware results
Emulated development results do not establish silicon acceptance. Physical-hardware validation remains ahead of us.

Chip development and validation at Memdance

Chip development and chip validation are both part of what we are building at Memdance. Development tools help engineers create, transform, and verify hardware designs. Validation tools help them examine how a machine behaves under controlled operations, system interactions, and sustained load.

Apex gives that validation effort its own execution environment, workloads, and approach to interpreting results. The broader direction is to support the engineering process from design through verification to validation on the machine, while preserving the distinct evidence each stage can provide.

Our next step for Apex is a successful physical boot followed by a bounded set of interpretable hardware results. Broader coverage and sustained stress will build on that foundation. This is how we intend to grow our chip-validation technology alongside our chip-development tools.