Switching hardware wallets? Migrate to Ledger safely in a few steps.

Learn more

Upgrade your digital life

Ledger Wallet: Free from compromise

Download now Learn more

Inference of Attacks by Bandits Against Zero-Knowledge Virtual Machines

Beginner
Ledger N3XT Research Competition

Automatically Hunting for Bugs That Let a Broken Computation Pass as a Valid Proof

Author
Ivan von Greiff
Blockchain Club
Statistical Cryptographic Attacks Group, TUM Blockchain Club
Track
Privacy
Date
September 2026
Student research published via the Ledger N3XT Research Competition. Findings are the author’s own. Ledger does not vouch for conclusions on advanced subject matter.
Abstract

Zero-knowledge virtual machines (zkVMs) let one computer run a program and prove to others that it delivered the correct results without requiring everyone to repeat the computation to verify. This guarantee depends on a constraint system correctly encoding the virtual machine’s program execution rules. If a required constraint is missing, an invalid execution may still produce a proof that the verifier accepts, creating a soundness bug. We study how to fuzz this security boundary more effectively. We introduce a new mutation surface that can alter internal zkVM data, and extract constraint system responses as feedback for enhanced scheduling of fuzzing decisions by our automated bug-finding framework. On RISC Zero, we find that the mutation surface determines which defect semantics are reachable, while constraint-guided scheduling improves accepted-proof yield when its feedback aligns with the reachable bug class. Constraint coverage is therefore useful as a conditional search signal rather than a universal proxy for soundness-bug discovery.

Contents

1. Introduction

Modern distributed systems increasingly rely on verifiable computation. One party performs a computation, while others check a compact proof instead of re-executing the computation. Zero-knowledge virtual machines (zkVMs) make this model programmable by allowing conventional programs to be compiled into a supported guest environment, executed, and accompanied by a proof that the virtual machine followed its rules and produced the claimed output.

This model is especially useful in blockchain systems, where nodes can validate blocks or smart-contract executions without repeating the full computation [1]. zkVMs also let developers prove general programs rather than designing a custom arithmetic circuit for each application, while the zero-knowledge property can hide private inputs without sacrificing verifiability. These benefits make implementation correctness critical, because if the VM’s constraints fail to capture its intended execution semantics, the verifier may accept a proof for an invalid computation.

Such a failure is a soundness bug: an execution that violates the intended virtual-machine semantics but nevertheless satisfies the implemented constraints and is accepted by the verifier. Soundness failures arise from underconstraints, where a required execution relation is absent or insufficiently enforced. The converse, a completeness bug, occurs when valid executions are rejected because the system is overconstrained. Both matter for zkVM correctness, but this work focuses on soundness. Our objective is therefore to deliberately corrupt the execution presented to the proving system and test whether an invalid computation can survive the remaining constraints and still reach verifier acceptance.

Program Inputs
↓
Guest Program
pub fn() { ... }
↓ Compiler
RISC-V Binary
01011 10110 01010
↓ RISC-V Executor
Execution Trace
{PCi, instri, reg1,i, …}i=1N
↓ Constraint System
Constraints
read(reg1) = prev write(reg1)?
instr = val(PC)? …
↓ Cryptographic Serializer
Proof
11001 01100 10111
↓
Verifier
Fig. 1. Pipeline from guest program to proof generation.

zkVM proof generation process. To understand where underconstraints can arise, consider how RISC Zero proves a program [2]. Figure 1 summarizes this pipeline. Rather than directly proving Rust or C++ semantics, the guest program is compiled into a 32-bit RISC-V executable. The proving problem is therefore reduced to showing that a sequence of machine instructions obeyed the expected RISC-V register, memory, arithmetic, and control-flow rules.

The RISC-V executor runs this binary instruction by instruction. Proving only the final output would not establish that the intermediate computation was valid, so the executor also records a step-by-step execution trace. Each trace row contains facts such as the program counter, executed instruction, register values, and relevant memory activity. For an instruction such as add x3, x1, x2, the trace records enough information to establish that the correct operands were read and that their sum was written to the intended destination. The completed trace therefore serves as the historical record from which correct execution is proven.

The constraint system turns this recorded history into a security claim by encoding the rules that a valid RISC-V trace must satisfy. For example, constraints can enforce that an instruction was decoded correctly, that its recorded source values came from the proper registers, and that its output was computed correctly. Other constraints enforce consistency across execution steps, such as requiring a later memory read to match the appropriate previous write. The prover constructs a cryptographic proof that the trace satisfies these implemented relations, and the verifier checks this proof without replaying every instruction. Soundness therefore depends on a critical correspondence, namely that satisfying the implemented constraints must imply that the intended RISC-V semantics were actually followed.

This pipeline also exposes an important distinction for testing. During execution, the executor maintains live machine state while it is still constructing the trace. Changing a register, instruction, or memory value at this point can therefore propagate into later execution. Post execution, the program has already finished and the trace represents a completed history, so modifying it changes the evidence supplied to the proving system without rerunning the computation. These two mutation points can create fundamentally different inconsistencies, making the distinction between live execution state and completed trace data central to this work.

Why ordinary testing is insufficient. Fuzzing searches for bugs by repeatedly generating or mutating tests and observing how the target responds [3]. Purely random testing can waste a finite budget revisiting behavior already exercised, so coverage-guided fuzzers instead reward tests that reach previously unseen program behavior and use a scheduler to allocate further effort toward productive regions. Coverage is not itself the security property being tested, but rather an observable proxy for novelty. An oracle separately determines whether an observed execution actually constitutes a bug. This separation between mutation, feedback, scheduling, and bug classification provides a useful starting point for zkVM fuzzing as well.

Input
→
Mutation
→
Execution
→
Crash /
Coverage
Feedback
←
Scheduler
←
Fig. 2. Feedback loop of a conventional coverage-guided fuzzer.

Applying this model to a zkVM is harder because the target is not merely one program, but an execution environment together with the proving system that claims its execution is valid. An invalid internal execution can produce the expected result and still be accepted, so crashes, path coverage, and output mismatches are insufficient soundness indicators. A zkVM soundness fuzzer must therefore decide what internal object to mutate, what constraint feedback to observe, how accepted outcomes are classified, and how a finite proof budget is allocated.

Prior work. The closest prior work is Arguzz, which combines metamorphic testing with fault injection and evaluated six RISC-V-based zkVMs, reporting eleven previously unknown soundness and completeness bugs [4]. For soundness testing, Arguzz simulates a malicious prover by perturbing registers, memory, program counters, fetched instructions, and other live state during execution. The corrupted state can propagate through subsequent instructions before Arguzz tests whether the resulting computation nevertheless produces an accepted proof. Arguzz therefore establishes that fault injection can uncover real zkVM soundness failures and provides the strongest direct baseline for our work.

More recently, zkvmBlast applies differential testing across multiple RISC-V zkVMs and a reference simulator to expose executor-correctness, completeness, and soundness failures [5]. This is complementary to our setting, as zkvmBlast searches for disagreements across implementations or against the RISC-V specification, whereas our framework deliberately corrupts internal proving artifacts and uses constraint-system failures to guide underconstraint search within one proving stack. Arguzz therefore remains our direct experimental baseline because it shares our malicious-prover fault-injection model.

Table I. Core zkVM-fuzzing design choices and the Arguzz baseline [4]
Choice Role Arguzz baseline / open gap
Mutation surface What internal object can be corrupted? During execution; cannot independently alter downstream trace metadata
Feedback What did the constraints report? No feedback
Oracle Did acceptance represent a semantic deviation? Verifier acceptance can include inert mutations
Scheduler How should finite proof budget be spent? Instruction-frequency balancing rather than constraint-aware allocation

Arguzz leaves two design dimensions particularly open. First, because its soundness mutations operate inside the executor, it cannot independently modify witness or trace metadata constructed only after the relevant execution event. Second, its scheduler balances injections across instruction types but receives no feedback describing how previous mutations interacted with the proving system. The usual analogy to coverage-guided fuzzing is therefore incomplete as the scheduler knows what it has attempted, but not which regions of the zkVM’s constraint system those attempts exercised. These gaps motivate our investigation of mutation surface and constraint-system feedback.

Effective zkVM fuzzing may therefore depend not only on generating adversarial mutations, but on matching where those mutations are applied with what the fuzzer can learn from their effects. We investigate three research questions:

  • RQ1. How does the mutation surface affect constraint coverage and the classes of soundness bugs a fuzzer can reach?
  • RQ2. When can constraint-system feedback improve mutation scheduling and bug-discovery effectiveness?
  • RQ3. Which framework components can remain zkVM-agnostic, and which must be adapted to the target machine and proving system?

To investigate these questions, we extend prior zkVM fuzzing along two dimensions. We introduce a post-execution mutation surface that exposes completed-trace metadata unavailable to executor-level fault injection, and instrument the proving pipeline to convert constraint failures into an online learning signal for adaptive mutation scheduling.

2. Methodology

Architecture and scope. Our framework forms a closed loop between an external fuzzing engine and a forked, instrumented RISC Zero target. At mutation step t, the scheduler selects an abstract mutation decision At, which is instantiated on the chosen mutation surface and executed through the proving pipeline. The target returns local and global constraint feedback (Γt, Zt), which the engine converts into campaign-wide coverage updates and a learning reward before choosing the next mutation. Mutation surface and scheduler can be varied independently while target invocation, feedback extraction, reward construction, and outcome recording remain shared, allowing us to separate the effects of where faults are introduced from how they are scheduled.

Fuzzer Engine — Python
MAB scheduler → mutation decision → reward / discovery bit
records → run.db (SQLite)
mutation decision At ↓ ↑ constraint feedback (Γt, Zt)
Target — forked, instrumented RISC Zero (risc0-host)
guest ELF + mutation-injection hook + constraint-failure emitters
Fig. 3. Feedback-guided fuzzing loop. The scheduler selects mutation At; the instrumented zkVM returns local and global constraint feedback (Γt, Zt) used to update coverage and the scheduling reward.

Our threat model treats the prover as potentially malicious or faulty: internal execution artifacts may be corrupted before the resulting computation is proven. A successful test produces a semantically invalid execution whose proof is nevertheless accepted. We therefore test the correspondence between intended RISC-V semantics and RISC Zero’s implemented constraints, rather than attacking the cryptographic proof primitives or zero-knowledge property themselves. Apart from mutation and diagnostic instrumentation, the proving protocol remains unchanged.

Mutation surfaces. Our first surface is the during-execution (DE) fault-injection interface introduced by Arguzz [4]. Its hooks operate while the RISC-V executor is fetching, decoding, evaluating, or committing an instruction and can perturb live state such as the program counter, registers, memory, fetched instruction word, or computed result. We classify these as DE mutations whenever guest execution continues after the fault, even when an individual hook occurs after one instruction’s semantics. Because subsequent instructions consume the corrupted state, a single DE mutation can propagate through later execution and alter multiple downstream trace rows.

Our second surface acts post-execution (PE), where the guest first executes normally and RISC Zero materializes the completed execution trace. Before proof construction proceeds, we intercept this trace and overwrite a selected recorded field without rerunning the executor. The scheduler chooses a coarse mutation class and semantic region, and the engine then inspects the trace, identifies compatible locations, and instantiates the concrete trace index and replacement value. Because execution has already ended, the modified value cannot propagate into later machine state, meaning PE mutation behaves as a localized edit to an otherwise completed history.

Moving the mutation point also exposes semantic objects that do not exist as independently writable state during execution. For example, Arguzz can change a fetched instruction word, but the executor subsequently decodes that same word and records matching opcode metadata. PE mutation can instead alter the completed opcode label while leaving the fetched word unchanged, creating a mismatch that executor-level mutation cannot independently express. Our implementation provides eleven PE mutation kinds spanning recorded register and memory values, instruction-decode metadata, and memory or lookup bookkeeping. We therefore treat DE and PE as distinct fault models and evaluate mutation surface as an experimental factor. The complete DE and PE mutation catalogs are reported in Appendix A, Table IV.

Constraint feedback and coverage. After each mutation, we instrument witness generation to observe how the constraint system responds. We first distinguish local constraints, which enforce relations within one instruction step or between neighboring steps. For every failed local relation, we record:

γ = (template, major, minor)    (1)

where template identifies the violated constraint and the opcode labels identify the RISC-V semantics under which it was instantiated. The opcode context matters because RISC Zero can reuse one constraint template across different instruction semantics. All local failures produced by mutation t are retained in Γt rather than terminating witness generation at the first failed relation.

Local failures do not capture relations spanning the full execution. RISC Zero also uses global memory and lookup arguments to enforce properties such as consistency between reads and writes. These checks aggregate many execution events using LogUp-style arguments [6], so an aggregate failure does not ordinarily reveal which individual events failed to cancel. We therefore instrument the contributing events and, when a global family is inconsistent, reconstruct the remaining uncanceled tuples. These tuples form the per-run global feedback set Zt.

The feedback (Γt, Zt) describes one mutation; campaign-level guidance requires remembering what has already been observed. We therefore maintain three growing sets: local contexts L, global contexts G, and structural mutation signatures S, whose concrete signature definition is given in Appendix A. Raw addresses and lookup indices are too granular for global coverage, so each uncanceled tuple is mapped through a compressed global-context function cgc(z) retaining properties such as global family, functional region, instruction/event class, and execution phase; the exact compression mapping is given in Appendix A. Coverage therefore means distinct observed contexts, not a percentage of a known finite constraint universe.

Coverage growth becomes the scheduler’s online reward. Let ℓnew(t) and gnew(t) denote previously unseen local and compressed global contexts produced by mutation t, and let snew(t) ∈ {0, 1} indicate a new structural mutation signature. We reduce this marginal growth to the binary event

Bt = 1[ℓnew(t) + gnew(t) + snew(t) > 0]    (2)

Thus, a mutation is rewarded for advancing a coverage frontier, but not rewarded more merely because it causes a large cascade of failures. Such cascades may indicate that the corrupted semantics are already heavily constrained rather than close to an underconstraint. We therefore treat novelty as an observable proxy for bug-discovery potential, and test the validity of that proxy empirically in Section 3.

Adaptive mutation scheduling. Given binary reward Bt, mutation selection becomes a sequential multi-armed bandit problem. Each arm is

a = (kind, zone)    (3)

where kind identifies the mutation operator and zone a coarse semantic region of the trace. The 19 semantic zones used in our implementation are listed in Appendix A, Table V. Exact trace positions and replacement values are intentionally excluded from the learning space. After an arm is chosen, trace inspection finds compatible locations and instantiates a concrete mutation. The scheduler therefore learns over semantically meaningful classes of decisions rather than every possible trace cell and value.

For each arm a, we estimate the probability that another pull expands at least one coverage frontier:

θa = P(Bt = 1 | At = a)    (4)

Because Bt is binary, we maintain θa ∼ Beta(αa, βa) with αa = βa = 1 initially. A discovery increments αa and a non-discovery increments βa. Thompson sampling [7] draws one candidate from each posterior and selects the largest. Arms with strong observed discovery rates are favored, while uncertain arms retain opportunities to be sampled due to their higher entropy.

Coverage novelty is, however, non-stationary because productive arms eventually exhaust the contexts they reach most easily. Early successes may therefore overstate an arm’s remaining value, while rarely sampled mutations may be abandoned too early. We combine Thompson sampling with an exploration floor, where at each decision, Et ∼ Bernoulli(φ0) is sampled. If activated, the least-pulled arm in the current exploration window is selected, otherwise Thompson sampling is used. We evaluate φ0 = 0.55, deliberately reserving substantial budget for persistent breadth.

Soundness-candidate oracle. Verifier acceptance alone does not imply a soundness failure because some mutations are semantically inert. We therefore replay and triage accepted outcomes using evidence from the mutated execution, separating audited no-ops from cases in which the represented execution genuinely changed or diagnostic global inconsistencies remain. The oracle prioritizes candidates for investigation rather than automatically declaring new vulnerabilities, and confirmation still requires deterministic replay and semantic inspection.

Experimental variants. We separate mutation-surface and scheduling effects along two axes. Mutations are applied either during execution (DE) through the Arguzz interface or post execution (PE) through our completed-trace interface. Scheduling is either feedback-agnostic Uniform, shorthand for Arguzz-style frequency balancing across eligible instructions or mutation classes [4], or the constraint-guided Bandit policy defined above. This yields four DE/PE × Uniform/Bandit single-surface configurations plus one hybrid DE+PE-Bandit configuration.

Table II. Mutation surfaces and scheduling policies of the evaluated variants
Variant Mutation Surface Scheduler
DE-Uniform (Arguzz)DE (baseline)Uniform (baseline)
PE-UniformPE (new)Uniform (baseline)
DE-BanditDE (baseline)Bandit (new)
PE-BanditPE (new)Bandit (new)
DE+PE-BanditDE+PE (hybrid)Bandit (new)

3. Evaluation

Our evaluation separates two aspects of fuzzing effectiveness: reachability, determined by which semantic inconsistencies a mutation surface can express, and search efficiency, determined by how the scheduler allocates a finite mutation budget within that space. We examine both through a constraint-coverage campaign, a reproduction of a known propagation-dependent RISC Zero soundness bug, and a controlled post-execution metadata underconstraint.

The coverage experiment evaluates our architectural variants on four guests stressing mixed arithmetic/branching, system/control-flow logic, memory-intensive execution, and SHA/accelerator behavior. Each (variant, guest) pair receives K = 5000 mutations, for 80,000 attempts total. Each bug race uses ten scheduler seeds and the same K = 5000 budget per seed. Because jobs ran on heterogeneous server hardware, comparisons use mutation budget rather than wall-clock runtime. Accepted-proof counts are compared only within each defect because their trigger frequencies differ.

3.1 Constraint Coverage and Mutation-Surface Behavior

Figure 4 shows a clear reversal between the two forms of coverage. PE variants reach substantially more distinct local contexts, while DE variants reach substantially more compressed global contexts. Configurations sharing the same mutation surface also cluster more closely than configurations sharing the same scheduler, and the same qualitative pattern appears across all four guest workloads. Mutation surface therefore appears to have the stronger first-order effect on the observed constraint profile, with scheduling producing a smaller within-surface difference.

Line charts of distinct local contexts and distinct compressed global contexts versus mutation index, comparing PE-Bandit, DE-Uniform (Arguzz), DE-Bandit, and DE+PE Bandit
Fig. 4. Representative local and global constraint coverage for the mixed-arithmetic and branching guest. The same qualitative DE/PE reversal appears across the other guest logic tested.

The reversal is consistent with the propagation semantics of the two surfaces. Let N denote the number of distinct global contexts broken by one mutation. Empirically, Pr(N > 10 | DE) ≈ 52% versus Pr(N > 10 | PE) ≈ 0.9%, while Pr(N = 0 | DE) ≈ 33% versus Pr(N = 0 | PE) ≈ 6.8%. DE faults can propagate through later execution, either remaining sufficiently self-consistent to bypass global checks or producing large cascades of downstream failures. PE point edits cannot propagate. Consistent with this distinction, every observed PE mutation kind except INSTR_TYPE_MOD (Appendix A, Table IV) triggered at least one global inconsistency. Ordinary execution-state values had already entered RISC Zero’s global bookkeeping, whereas the derived instruction-type label had not. These results also show why coverage is only a proxy: breaking many constraints can indicate strong enforcement rather than proximity to an underconstraint.

3.2 Bug-Discovery Races

Historical during-execution bug. We reproduce CVE-2025-52484, a previously disclosed and patched RISC Zero soundness bug identified by Arguzz [4], [8], [9]. A three-register instruction can be mutated so that both source operands reference the same register, for example changing remu a3, a0, a1 to remu a3, a0, a0. Because execution continues from the corrupted instruction, the duplicated reads enter the downstream machine history. In the vulnerable implementation, their memory-multiset contributions can self-cancel without enforcing the intended source-register relation, allowing the invalid execution to remain provable. We test only the patched bug’s historical vulnerable revision.

Post-execution metadata underconstraint. We then deliberately remove the constraint linking a fetched instruction word to the major/minor instruction-type label recorded in the completed trace. This is a manufactured underconstraint, not an upstream RISC Zero vulnerability. PE mutation can directly create the resulting mismatch by changing the completed label while leaving the instruction word unchanged. DE mutation cannot independently do so because the label is derived downstream, keeping both representations mutually consistent. The experiment therefore reverses the previous surface requirement. The historical defect requires propagation through live execution, whereas this defect requires downstream trace fields to be decoupled only after execution has completed.

Table III. Accepted proofs across ten seeds in both bug races. Counts are comparable within, not across, defect columns because the two defects have different trigger frequencies.
Variant DE bug PE metadata bug
DE-Uniform90
DE-Bandit90
DE+PE-Bandit3478
PE-Uniform0588
PE-Bandit0755

Table III shows the expected reversal. For the historical bug, DE-Uniform and DE-Bandit each produce nine accepted proofs, while both PE variants produce none. DE-Bandit’s greater global coverage therefore does not improve discovery of this local underconstraint; PE instead fails because a post-hoc trace edit cannot recreate the propagated execution history. For the metadata defect, both DE variants produce zero accepted proofs, while PE-Bandit produces 755 versus 588 for PE-Uniform, an increase of approximately 28%. Here, local-context exploration aligns with an instruction-level metadata (local) relation. Constraint feedback therefore improves yield only after the mutation surface makes the relevant semantic inconsistency expressible.

3.3 Synthesis: Surface-Aligned Feedback

Together, the experiments establish that reachability precedes search efficiency. Mutation surface first determines whether the target inconsistency can be expressed; constraint feedback then guides budget allocation within that space. DE+PE-Bandit reaches both defect classes but underperforms the best surface-specific configuration because its fixed budget spans a larger action space and its exploration floor continues sampling under-used decisions. Effective fuzzing therefore requires alignment among mutation surface, feedback signal, scheduler, and target semantics.

Ledger Lens

Settlement Trust Has Two Different Sources

Ledger’s “Revenge of the Atoms” thesis argues that critical digital trust should be anchored in hardware rather than left entirely to mutable software [10]. Ledger’s broader security roadmap applies this principle to secrets and authorization, using hardware as an enforceable root of trust even when surrounding software is compromised [11]. Zero-knowledge proofs address a different question i.e. whether a computation or state transition satisfied specified rules without requiring every observer to repeat it.

That distinction matters because a valid proof guarantees satisfaction of the implemented constraints, not that engineers implemented every constraint that was required. A soundness bug breaks this correspondence, where an invalid execution satisfies the rules that were written while violating one that should have been. The verifier then behaves correctly relative to an incorrect specification. Confidence in this boundary therefore requires adversarial testing of what the proving system actually enforces, not merely confidence in the cryptography.

This creates a direct connection between hardware-rooted trust and zkVM soundness without treating them as substitutes. A secure element can establish that an action was authorized by a protected key, but it cannot establish that an upstream rollup state or proof-verified computation was semantically correct. Conversely, a sound proving system cannot protect a compromised signing key or guarantee user authorization. Hardware anchors authority, while sound verification anchors computational validity.

Our results add a final qualification: adversarial testing is itself surface-dependent. Different points in a proving pipeline expose different semantic objects and failure classes. The broader Ledger-relevant lesson is to identify each trust boundary, state its assumptions explicitly, and test the mechanism enforcing them rather than treating either hardware or cryptography as a universal security guarantee.

4. Conclusion

RQ1 shows that mutation surface changes both observed constraint behavior and reachable bug classes, while no tested surface dominates across defects. For RQ2, constraint failures can guide scheduling when rewarded novelty aligns with reachable defect semantics, but coverage is not a universal predictor of bugs. RQ3 is answered only architecturally. The bandit and reward abstraction are the most reusable components, while semantic mappings are partially portable and mutation hooks, trace representations, constraint instrumentation, and oracle logic remain target-specific.

We evaluate one zkVM, one patched historical soundness bug, and one manufactured metadata underconstraint. The results therefore do not establish general superiority of PE over DE, Bandit over Uniform, or local over global coverage, while cross-zkVM portability also remains to be demonstrated empirically.

Future work should therefore build mutation catalogs specifically around the semantic objects exposed at different proving-pipeline surfaces, rather than applying one generic fault model everywhere. Richer rewards should distinguish local from global novelty and sparse, isolated failures from large rejection cascades, bringing the learning signal closer to the true objective of finding weakly constrained semantics. The hybrid results also motivate hierarchical scheduling that first allocates budget between mutation surfaces and then learns which mutations to prioritize within each surface. Finally, these ideas should be tested against a larger corpus of natural soundness bugs and across multiple zkVM architectures.

Appendix

Appendix A. Supplementary Implementation Details

Mutation catalogs. The evaluated mutation surfaces operate on different representations of execution. DE mutations perturb live machine state through the Arguzz fault-injection interface [4], whereas PE mutations edit the completed trace after guest execution has ended. The complete catalogs are shown in Table IV. Mutation identifiers retain their implementation names across the two surfaces.

Table IV(a). During-execution (DE) fault-injection catalog. Execution continues after every mutation.
Kind What it perturbs Phase Effect
PRE_EXEC_PC_MODProgram counterPreExecution begins the instruction from a modified address.
INSTR_WORD_MODFetched instruction wordPreA different instruction word is decoded and executed.
PRE_EXEC_MEM_MODMemory wordPreMemory state is corrupted before the instruction consumes it.
PRE_EXEC_REG_MODRegister valuePreA register is corrupted before its value is read.
BR_NEG_CONDBranch-taken conditionPreThe branch decision is inverted.
COMP_OUT_MODComputed resultPostA corrupted result is written to the destination register.
LOAD_VAL_MODLoaded valuePostA load returns a corrupted value.
STORE_OUT_MODStored valuePostA corrupted value is written to memory.
POST_EXEC_PC_MODNext program counterPostThe next instruction fetch is redirected.
POST_EXEC_MEM_MODMemory wordPostMemory is corrupted after the current instruction.
POST_EXEC_REG_MODRegister after the instruction stepPostA register is corrupted after the current instruction.
Table IV(b). Post-execution (PE) mutation catalog. All mutations directly edit a completed execution trace without re-execution or forward propagation.
Kind Trace cell perturbed Constraint targeted Effect
COMP_OUT_MODDestination-register write valueRegister-write consistencyRecords an incorrect result for the destination register.
LOAD_VAL_MODDestination-register write valueMemory-read consistencyRecords an incorrect value returned by a load.
STORE_OUT_MODMemory-write valueMemory-write consistencyRecords an incorrect value written to memory.
MEM_VAL_MODMemory transaction valueMemory-read/write consistencyChanges the value recorded for a memory transaction.
PRE_EXEC_REG_MODRegister transaction wordRead/write-chain consistencyMakes a recorded read disagree with the preceding write.
INSTR_WORD_MOD_FULLFetched word and prev_wordDecode / opcode verificationReplaces the complete recorded instruction word.
INSTR_WORD_MOD_SURFetched word and prev_wordDecode sub-checkSurgically changes one instruction field, such as opcode, register index, or function bits.
INSTR_TYPE_MODCycle major/minorInstruction type versus fetched wordMakes the claimed instruction type disagree with the fetched instruction word.
TXN_PREV_WORD_MODtxn.prev_wordRead/permutation chainBreaks the relationship to the previously recorded word.
TXN_PREV_CYCLE_MODtxn.prev_cyclePermutation temporal orderingBreaks the ordering of same-address memory events.
CYCLE_DIFF_COUNT_MODcycle.diff_count[i]Cycle-difference lookup / CycleArgIntroduces an inconsistency in lookup multiplicity.

Global-feedback compression. Raw global memory and lookup failures have a very large identity space, so we map them into coarser semantic contexts. For memory and lookup/range failures, respectively,

cgc(zmem, a) = (F, ρ(addr), β(addr), τ(kind), ψ(zone))    (5) cgc(zlook, a) = (F, β(index), kind, oc(major))    (6)

Here F identifies the global-constraint family, ρ maps an address to a functional memory region, and β(x) = ⌊log₂ max(x, 1)⌋ provides a logarithmic magnitude bucket. Transaction role τ distinguishes read, write, ifetch, register, prev_word, and prev_cycle; lifecycle phase ψ distinguishes normal, ecall, mret, halt, and boundary; and oc groups instruction semantics into alu, mul, div, mem, branch_or_ctrl, poseidon, sha, or other.

Scheduler semantic mapping. The scheduler learns over the coarse semantic decision

a = (kind, zone)    (7)

rather than individual trace indices or replacement values. Here kind selects a mutation and zone identifies the semantic region in which it is instantiated; the 19 evaluated zones are listed in Table V. Structural coverage further distinguishes concrete realizations through

σ = (kind, zone, substrategy)    (8)

For example, INSTR_WORD_MOD_SUR can alter different instruction fields while remaining the same scheduler arm. Tracking σ allows these realizations to contribute separately to structural novelty.

Table V. Semantic zones used by the mutation scheduler
Class Zone Meaning
Boundarystep0First execution step, corresponding to boot or initialization.
Boundarylast_stepFinal execution step, corresponding to termination or halt.
Boundarypre_ecallStep containing an ECALL cycle, corresponding to system-call entry.
Boundarypost_ecallFirst user-PC decode step shortly after an ECALL, corresponding to return to user execution.
Boundarypre_mretMachine-lifecycle region immediately before MRET.
Boundarypost_mretMachine-lifecycle region immediately after MRET.
Boundarypre_haltRegion immediately before halt.
Boundarypost_haltRegion immediately after halt.
Corecore_arithmeticArithmetic and logical operations such as ADD, SUB, XOR, OR, AND, and SLT.
Corecore_mulMultiply instructions.
Corecore_divDivision and remainder instructions, including DIV, DIVU, REM, and REMU.
Corecore_shrRight-shift instructions, including SRL, SRA, SRLI, and SRAI.
Corecore_memory_loadLoad instructions.
Corecore_memory_storeStore instructions.
Corecore_branchBranch and jump instructions.
Corecore_shaSHA gadget execution.
Corecore_poseidonPoseidon gadget execution.
Corecore_otherFallback for remaining core decode classes.
Kernelkernel_otherPrimary decode cycle at a kernel program counter.
References
  1. R. Lavin, X. Liu, H. Mohanty, L. Norman, G. Zaarour, and B. Krishnamachari, “A Survey on the Applications of Zero-Knowledge Proofs,” 2024. arxiv.org/abs/2408.00243
  2. RISC Zero, “RISC Zero zkVM,” GitHub repository, 2026, accessed 2026-09-03. github.com/risc0/risc0
  3. V. J. M. Manès, H. Han, C. Han, S. K. Cha, M. Egele, E. J. Schwartz, and M. Woo, “The art, science, and engineering of fuzzing: A survey,” IEEE Transactions on Software Engineering, vol. 47, no. 11, pp. 2312–2331, 2021. doi.org/10.1109/TSE.2019.2946563
  4. C. Hochrainer, V. Wüstholz, and M. Christakis, “Arguzz: Testing zkVMs for soundness and completeness bugs,” Baltimore, MD, Aug. 2026. usenix.org/conference/usenixsecurity26/presentation/hochrainer
  5. S. Chaliasos, M. Ochoa, and V. Thakore, “Introducing zkvmBlast: Differential Fuzzing for Ethereum’s zkVMs,” ZK/SEC Quarterly, Aug. 2026, zkSecurity research blog, accessed 2026-09-03. blog.zksecurity.xyz/posts/zkvmblast
  6. U. Haböck, “Multivariate Lookups Based on Logarithmic Derivatives,” Cryptology ePrint Archive, Paper 2022/1530, 2022, version dated November 4, 2022. eprint.iacr.org/2022/1530
  7. W. R. Thompson, “On the likelihood that one unknown probability exceeds another in view of the evidence of two samples,” Biometrika, vol. 25, pp. 285–294, 1933. semanticscholar.org/CorpusID:120462794
  8. flaub, “ZKVM-1392: Disallow Memory I/O to Same Address in the Same Memory Cycle,” GitHub pull request, risc0/risc0, PR #3181, May 2025, merged May 23, 2025. github.com/risc0/risc0/pull/3181
  9. jbruestle, “ZIR-366: Fix to Remove Extra Register Read When Both Source Registers Are the Same,” GitHub pull request, risc0/zirgen, PR #238, May 2025, merged May 23, 2025. github.com/risc0/zirgen/pull/238
  10. P. Gauthier, “Revenge of the Atoms,” Ledger, Mar. 2026, ledger blog, accessed 2026-09-03. ledger.com/blog-revenge-atoms
  11. I. C. Rogers, “Securing Your Agents With A Hardware Root of Trust: Ledger’s 2026 AI Security Roadmap,” Ledger, Apr. 2026, ledger blog, accessed 2026-09-03. ledger.com/blog-2026-ai-security-roadmap
Originality Statement

I confirm that this submission is an original research paper produced for the Ledger N3XT Research Competition (Privacy Track). The original contributions, implementation extensions, empirical analyses, and conclusions presented in this manuscript are my own work. Prior methods, external codebases, academic literature, and other source material are clearly distinguished from these contributions and cited where applicable.

This manuscript is free from plagiarism and has not been published or submitted concurrently to any other conference, journal, or competition. I assume full responsibility for the research, methodology, and conclusions presented.

Ivan von Greiff  ·  September 2026


Stay in touch

Announcements can be found in our blog. Press contact:
[email protected]

Subscribe to our
newsletter

New coins supported, blog updates and exclusive offers directly in your inbox


Your email address will only be used to send you our newsletter, as well as updates and offers. You can unsubscribe at any time using the link included in the newsletter. Learn more about how we manage your data and your rights.

Own your crypto future

Stay informed with security tips, updates, and exclusive offers from Ledger

Your email address will only be used to send you our newsletter, as well as updates and offers. You can unsubscribe at any time. Learn more

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.