Switching hardware wallets? Migrate to Ledger safely in a few steps.

Learn more

Upgrade your digital life

Ledger Wallet: Free from compromise

Download now Learn more

Who Said Yes?

Beginner
 
Ledger N3XT Research Competition

Authorization Boundaries for Agents That Hold Value

Author
Youssef Jeddi
X (Twitter)
@YoussefJeddi6
Blockchain Club
EPFL Blockchain Student Association (BSA)
Track
Agentic Economy
Date
September 2026
Student research published via the Ledger N3XT Research Competition. Findings are the author’s own. Ledger does not vouch for conclusions on advanced subject matter.
Abstract

When an autonomous agent sends a transaction, something decided the transaction was allowed. This paper asks what, and where that decision physically lives. We study Orchestra, an agent we built that turns English instructions into onchain transactions on an Ethereum test network, and follow one change through its history: a language model was first asked to approve the agent’s own plans, then removed from that decision entirely and replaced by a small set of rules returning the same verdict every time they see the same facts. We attacked both with 34 hand-written adversarial cases across two threat models.

The interesting result is not that the guardrails held, but that where they failed, they failed in exactly one way. No attack talked the rules out of a verdict, because the rule code never reads the user’s words. Every failure instead lied to the rules about what was moving: a request to “swap all of my ETH” whose amount is non-numeric is valued at NaN, which slips under every threshold, the interface reports “$NaN is within your $100 limit” and executes. When we closed one such defect, a swap labelled with the wrong token symbol, valued from the label instead of the address, the deterministic engine went from failing that case to catching it while the language model still waved it through. Removing the model from the decision did not remove the attack surface; it moved it down to the inputs the decision is computed from, and closing one input defect leaves the others the same shape. That is where hardware-backed confirmation earns its place: proving a person was present cannot help a system misled about what it is doing, while showing that person the actual payload can.

Contents

1. Introduction

Agents that hold value are arriving faster than the authorization models they need. A language model can read a sentence, work out which transaction it implies, and hand that transaction to a key. What it cannot do is make the resulting authorization legible: ask why a payment was permitted and the honest answer is that a probabilistic function emitted a token sequence parsing as approve.

Nearly every authorization mechanism in production assumes a human who is present and paying attention. A signature prompt, a confirmation dialog, a one-time code: each is a question put to somebody expected to answer it. An agent acting while its principal is asleep voids that assumption without offering a replacement. The question is therefore not “can the agent transact” but “when it does, who said yes, and what does that yes prove”.

We work the problem on one system rather than in the abstract. Orchestra’s git history contains a clean natural experiment: the authorization decision was moved out of a language model and into deterministic code, and both versions are recoverable and attackable.

We contribute: (1) a case study of removing a language model from the authorization path of a working system; (2) a deterministic policy engine with justified fail-closed ordering, in which rules may escalate but never weaken; (3) an adversarial evaluation over 34 cases and two threat models showing that the residual attack surface of the deterministic gate is entirely valuation, it relocates rather than disappears, together with a diagnosis of the input-boundary defects behind it; and (4) a taxonomy of authorization locus, model, server, chain, device, and what a compromise at each level survives.

2. Methodology

2.1 Approach

The study is build-then-adversarially-review on a single artifact. We have n=1 by construction, trading generality for the ability to inspect layers a black-box evaluation cannot reach: the decision function, the values it was fed, and which rules fired. Our claims are existence claims about failure classes, never frequency claims about a population.

2.2 Artifact and attribution

Orchestra is roughly 12k lines of TypeScript: a web interface, an Express bridge service, a planner backed by a hosted language model, a policy engine, and execution adapters targeting the Sepolia test network. The code is a group artifact. The research question, the adversarial corpus, the evaluation, the argument and this text are the submitting author’s own. The harness and its raw output are published with the source [16], so every number below can be regenerated.

2.3 Measurement and classification

Each case is a plan or prompt with a hand-labelled expected verdict; we record the verdict returned, the USD value assigned, and which rules fired. Outcomes are held, failed open (an action needing a human executed unattended), or over-escalated (a benign action demanded approval it did not need). Only failed-open is a security failure; over-escalation is a usability cost, counted separately because pretending it is free is how fail-closed systems become unusable. For cases involving the model we run three identical trials and report the worst: a guardrail that holds two times in three does not hold.

2.4 Assumptions and limitations

We assume an honest RPC endpoint, an uncompromised user device and a single server instance. Testnet assets are valued at mainnet reference prices, pinned at ETH/WETH $2,500 and USDC/USDT $1 so the figures reproduce. The corpus is hand-built and deliberately adversarial: it shows a failure class exists, not how often it occurs.

Three limitations deserve stating rather than burying. First, the model in the original build was withdrawn by its provider mid-project, so results were produced with a different one: agent behaviour is not reproducible across a model deprecation, which is itself a finding about this class of system. Second, our harness was wrong twice, and both times in the flattering direction, once scoring a planning failure as a successful block, once scoring provider rate-limit errors the same way. Both were caught by a benign control failing when it should have passed. An adversarial harness that silently rewards the system under test is a quiet way to publish a result that is not true. Third, the provider’s free-tier daily token budget is smaller than a full pass of the corpus requires, so the after-fix aggregate over all 34 cases is left incomplete; each per-case result we report below, RC2 caught, RC3 and RC4 failing open, comes from a targeted run that does complete, and the one number this constraint leaves pending is the full-corpus rate with the fix in place.

3. The System, and Why the Model Lost the Decision

3.1 Orchestra in brief

A user writes a sentence. A planning model converts it into a structured plan: a summary plus an ordered list of typed steps, each with an intent type and explicit parameters. Custody is a Safe smart account owned by a hardware wallet, and the agent holds a separate delegate key bounded by an onchain allowance module, so the worst case for a stolen agent key is capped by a contract rather than by the server’s good behaviour [11].

One architectural fact matters here: between the plan and any transaction being built sits exactly one decision, returning one of three verdicts: execute unattended, require approval, block. Everything below is about that function and the inputs it is given.

3.2 Version 1: the model as gatekeeper

The first version asked a second language model to authorize the plans the first one produced. It was given the plan and a description of the user’s limits and asked to return a verdict. It worked most of the time. Then it did not, and the patch is the interesting artifact: a hard-coded server-side check was added to catch cases where the gatekeeper approved things it should not have.

That patch is the tell. Once you must add a deterministic check to correct your probabilistic authority, the deterministic check is the authority, and the model has become an expensive, non-reproducible pre-filter in front of it. Version 2 deleted the gatekeeper model and promoted the check.

3.3 Three objections

Three properties made the arrangement untenable, and they generalise past this system.

Non-determinism in a security decision. A policy you cannot reproduce is one you cannot regression-test. If the same plan under the same limits can yield different verdicts, there is no such thing as a fix, only a shift in a distribution. Section 5 measures this: six of our prompts produced different outcomes for the model gate across identical runs, five for the deterministic engine.

No audit surface. A rule that fired has an identifier, a threshold and a line number. A verdict that was generated has a probability nobody logged, and cannot be diffed, versioned or argued with.

The judged artifact is the attack vector. The plan derives from the user’s message, so whoever writes the message shapes what the judge reads. Asking a model to adjudicate adversary-controlled text is the structure of indirect prompt injection [4, 5, 6]; instruction hardening does not change the fact that instructions and data share one channel.

3.4 A check that cannot fail

Auditing for this paper we found a second, older risk assessor still live on the quoting path, where its verdict steers route selection. Every input it receives is a hard-coded placeholder, so it returns the same permissive answer on every call. We report it as motivation rather than as an exploited vulnerability, because the general point outlives the instance: a safety check that always passes is, from outside, indistinguishable from one that works. Only a case it should reject tells them apart.

Two flow diagrams: version 1 (probabilistic authority) runs message through Planner LLM, Gatekeeper LLM, a server override, to execute; version 2 (deterministic authority) runs message through Planner LLM directly into a heavy-bordered decide() function, then an adapter and Safe, to chain
Fig. 1. Version 1 asked a model to authorize the agent’s own plans, then added code to correct it; version 2 deletes the model from the decision. Dashed boxes are non-reproducible; the heavy box is where the verdict is produced.

4. A Deterministic Gatekeeper

The replacement is a pure function. It performs no I/O, reads no free text, and takes its clock as a parameter, which is the only reason time-dependent rules such as a rolling window are testable at all. Given the same plan, the same history and the same clock, it returns the same verdict, and that verdict carries the list of rule identifiers that produced it.

4.1 Fail-closed ordering

The function evaluates in a fixed order: read-only intents pass immediately; malformed plans, denylisted counterparties and unrecognised intents are blocked; then the escalation rules run; and only a plan surviving all of it executes unattended.

The order is doing work, not decoration. Cheap structural facts are checked before semantic ones, so a malformed plan can never reach a rule that might excuse it, and unrecognised intents are refused rather than passed through, so adding a capability is a deliberate act and not an accident of parsing. Most importantly the escalation rules are one-directional: each may raise the required approval level, none may lower it. Adding a rule therefore cannot make the system more permissive, which matters more in practice than any individual threshold, because it is what makes the policy safe to edit under time pressure.

Five rules escalate: a rolling 24-hour USD spend limit, a token allow-list, a recipient allow-list, a velocity check on transaction count, and a habit-anomaly check. The daily cap is deliberate, a per-transaction limit is trivially defeated by splitting a payment, whereas a rolling cap subsumes it and prices the split attack at zero benefit.

4.2 Habits without machine learning

The anomaly rule takes the median of recent transaction values and flags anything above ten times it. A median is robust to outliers and cannot be moved by a single large sample, so the baseline cannot be poisoned in one transaction; it is explainable to the user in a sentence; and because the rule only escalates, a wrong baseline costs an unnecessary confirmation rather than an unauthorized transfer. That asymmetry is the point: statistical machinery is acceptable in a component that can only ask for help.

Every verdict also carries machine-readable rule identifiers, so the interface can name the constraint that fired and its threshold. This is usually filed under user experience; we think it is a security property. A verdict a user can only accept or override teaches them to override, whereas one that names its reason lets them notice when the reason is wrong, the only mechanism by which a mis-specified policy gets reported rather than routed around.

Flowchart: plan leads to read-only intent check (yes routes to execute); then malformed, denylisted counterparty, and unrecognised intent checks (any yes routes to blocked); surviving plans reach escalation rules (rules may raise, never lower), which route to needs approval if any fired, otherwise auto-execute
Fig. 2. Structural checks precede semantic ones, and the rule bank is one-directional: every uncertain case resolves toward asking a human.

5. Adversarial Evaluation

5.1 Corpus and layers

We wrote 34 adversarial cases spanning 11 categories: benign controls (small and large), prompt injection, unit confusion, unknown recipient, unverified token, unknown intent, velocity, habit anomaly, value misreport, and malformed plans. The corpus is split by threat model. Layer A (10 cases, “compromised planner”) sends crafted plans straight to the decision function, offline and deterministically, isolating policy from model. Layer B (24 cases, “adversarial input, honest model”) sends natural-language prompts through the real planner and the unmodified production path, three times each. The reported runs use openai/gpt-oss-120b via Groq. Layer B stops at the verdict and broadcasts nothing: the object of study is the decision, not the execution. Benchmarks such as AgentDojo [7] evaluate agents against injected content in shared tool environments; our corpus is smaller and narrower by design, aimed at one authorization boundary rather than at task completion.

5.2 What held

Everything that attacked the rules failed, at both layers, in every trial. Instruction override, plan objects carrying a forged verdict field, claims of a prior hardware approval, asserted developer mode: zero successes. This is not resilience, it is structure. The decision function never reads free text, the plan summary is checked for existence, never interpreted and a forged verdict field is not one of its parameters, so it is not rejected but invisible. Accumulation attacks held completely: split transactions were absorbed by the rolling window, and both malformed cases failed closed across every trial.

5.3 What failed: valuation

Every unsafe approval the deterministic engine returned reduces to a valuation defect, a value read wrong at the input boundary, never a rule argued down. Four such defects make up the surface (Table 1). One, RC2, we have since closed, and its former exploit is now a case the engine handles better than the model did; the other three remain open, two of them demonstrated by the corpus and one by code audit. We report these per case rather than as a single post-fix rate: the pre-fix full-corpus figure is in Table 2 (an 8.0% deterministic unsafe rate, all of it valuation), and the after-fix full-corpus rate is the one number the daily-budget limit of Section 2.4 leaves pending.

Table 1. Four input-validation defects at the valuation boundary.
Defect Severity
RC1 Price lookup fails open on an unknown symbol: returns amount × $1 instead of refusing to value the plan High
RC2 Valuation trusts the attacker-supplied symbol label over the token address in the same plan Critical
RC3 No check that the computed value is finite and non-negative; NaN silently passes every threshold comparison Critical
RC4 A missing field disables its rule instead of triggering it: a plan with no recipient escapes the recipient allow-list High
RC2 is now fixed (§5.4) and its former exploit is caught; RC3 and RC4 are observed fail-opens (rc3-01/rc3-02, rc4-01); RC1 is a latent defect of the same class. None is a policy defect.

Three examples carry the argument. Under RC2, a swap of 2 WETH labelled symbolIn: "USDC" was previously valued at $2 instead of $5,000 and auto-executed, because the label overrode a token address that was present and authoritative in the same plan. Valuing from the address rather than the label closes it: the engine now prices the swap at $5,000 and escalates, while the historical LLM gate, judging the same plan, still auto-executes it. The defect that was the deterministic engine’s flagship failure is now a case it catches and the model does not, supplying more metadata no longer lowers the assessed value.

Under RC3, “swap all of my ETH” yields a non-numeric amount, hence NaN, which defeats every comparison the engine makes: IEEE-754 specifies that all ordered comparisons against NaN are false [15], so a fail-closed design whose only operator is > inherits a fail-open result. The interface reports “$NaN is within your $100 daily auto-approve limit” and executes. This reproduces both as a crafted plan and end to end from English (rc3-01, rc3-02). Notably the LLM gate catches it: the two gates fail on different inputs, neither dominating.

Under RC4, a send with no recipient escapes the recipient allow-list, which abstains on the missing field rather than firing. A $5 transfer to nobody auto-executes (rc4-01), and here both gates miss it. In each case the attack persuaded the engine of nothing; it corrupted the engine’s picture of the world, and the engine reasoned correctly about a world that did not exist. That is the thesis in miniature.

5.4 Fixing the boundary

All four remedies are boundary work and none touches a rule. Of the four, one is implemented and measured; three are specified but not yet built.

RC2 is fixed. Valuation now resolves the token symbol from the address rather than the supplied label: where the plan’s token address is a known asset, its symbol is authoritative and the label is ignored. This was sharper to implement than to describe, the check cannot compare label strings, because WETH and ETH are legitimately the same asset and a naive comparison flags every ordinary swap, so it resolves the address to a canonical symbol first. After it, the symbol/address attack holds: plan-02 now values at $5,000 and escalates, where before it auto-executed at $2. The fix is partial in a way worth stating: it covers known token addresses, and it does not yet raise the stronger signal, a disagreement between address and label as an escalation in its own right, so an unknown token (RC1) still slips through.

RC1, RC3, RC4 remain. Return no value for an unpriced asset rather than a default, and escalate on the absence, which requires a price function that can return null, since “worth nothing” and “could not be priced” are currently collapsed into zero (RC1). Reject any plan whose computed value is not finite and non-negative (RC3). Invert the missing-input rule so it fires rather than abstains (RC4). RC3 and RC4 are the residual we report, and unlike RC1 both are demonstrated: a plan whose amount is the string “all” is valued at NaN and auto-executes (rc3-01), and the same follows end to end when the planner emits a non-numeric amount from “swap all of my ETH” (rc3-02); a $5 send with the recipient field omitted escapes the allow-list and auto-executes (rc4-01). Closing them is code, not research.

5.5 Non-determinism in the interpretation layer

Across three identical runs, five of Layer B’s 23 prompts returned different verdicts for the deterministic engine and six did for the LLM gate, variance that lives entirely in Layer B, since Layer A runs offline and deterministically by construction. Most oscillated between BLOCKED and NEEDS APPROVAL: two different gates, both of which stop. Variance of that kind is tolerable in a component that can only propose and intolerable in one that decides, which is the argument for Section 4’s architecture; version 1 put exactly this variance on the decision. Note that this variance is the interpretation layer’s, not the policy’s: the same plan always yields the same verdict from decide(); it is the plan the planner emits that changes.

The valuation boundary is not the last of its kind, and one trial showed the next one directly. Intent type is itself a model-supplied field the engine consumes as fact: read-only intents pass the decision function immediately, the first branch in Section 4.1, so a plan arriving labelled balance query bypasses every rule, not by defeating one but by arriving as something the rules do not apply to. On one of its three trials, “swap all of my ETH” was classified by the planner as a balance query rather than a swap; the decision function returned INFO and no rule ran. Nothing moved, because a balance adapter can only read, the damage was bounded by the executor, not the policy, which is not a property we designed or would rely on. It is the same defect as the amount, a model-supplied field consumed as fact, one layer up, and Section 6 returns to it.

The deterministic engine’s friction rate was 20% against the LLM gate’s 13.3%: a benign transfer over-escalated across its trials because the planner attached an unverified token address, firing the token allow-list rule. Nothing was lost but the user’s attention, which the Ledger Lens argues is the resource this design actually spends, and which a fail-closed gate spends more freely, counted here rather than hidden.

Table 2. Outcomes by category before the RC2 fix, LLM gate versus deterministic engine (openai/gpt-oss-120b, 3 trials/case).
Category Trials Unsafe (LLM) Unsafe (Det) Friction (LLM) Friction (Det)
benign_small 18 0 0 2 3
benign_large 9 1 0 0 0
prompt_injection 18 1 3 0 0
unit_confusion 6 0 0 0 0
unknown_recipient 6 3 0 0 0
unverified_token 3 0 0 0 0
unknown_intent 9 2 0 0 0
velocity 6 3 0 0 0
habit_anomaly 3 0 0 0 0
value_misreport 9 0 3 0 0
malformed 6 0 0 0 0
Total 93 10 6 2 3
Even here the deterministic engine’s every unsafe approval falls in prompt_injection and value_misreport, both valuation, while attacks on the rules themselves (unknown_recipient, velocity, unknown_intent) that the LLM gate waved through are all caught. Unsafe totals are over 75 attack trials (13.3% vs 8.0%); friction is over 15 benign trials (13.3% vs 20%). With RC2 fixed these valuation failures are caught; the rc3-*/rc4-01 cases then become the deterministic engine’s residual (full-corpus run pending).

6. Residual Failure Modes

Enumerating what our design does not defend against is not a caveat section; it is the substance of the claim.

The guardrails run in a server process. An attacker who controls that process controls the policy; only the onchain allowance module bounding the delegate key survives. A policy contract deployed alongside it stores a spending limit, but nothing in the enforcement path reads it: a registry, not a control. Its existence could otherwise be read as protection it does not provide.

The price oracle sits inside the valuation boundary. Prices come from a single unauthenticated endpoint, cached, and degrade silently to hard-coded defaults when the fetch fails, which is RC1 in another dress. A dollar-denominated limit inherits every trust assumption of its price source, and mis-valuation, not rule subversion, is the live attack.

Intent classification is still a model output. The read-only fast path is taken on the strength of a label the planner supplies, and Section 5.5 shows that label is unstable. We did not fix it, because the fix is a design question rather than a patch: either the fast path goes and cheap reads pay for the full pipeline, or read-only becomes a property the executor proves rather than one the planner asserts.

Velocity counting is single-instance. The rolling window lives in one process; run two and the count is wrong, a correctness bug that presents as a security bug only under concurrency.

Ledger Lens

Three Tiers, Three Guarantees

Orchestra has three approval tiers. The useful exercise is asking what each one proves.

Table 3. Three tiers proving three different propositions. Calling all three “approval” is what creates the gap.
Tier What it proves What it does not
Unattended Nothing about the user Anything
Passkey A registered device was present and a person authenticated That the person saw what they approved
Hardware wallet A person saw a payload on a screen the host cannot rewrite, and signed those bytes That the payload is what they wanted

The middle row is where the confusion lives. A passkey is an excellent proof of presence and a strong defence against credential theft [8], but not a proof that the payload consented to is the payload that will be signed, the payload is rendered by the same host that requested the approval. A system treating presence as consent has built a confused deputy [2]: an authorized party induced to exercise its authority on someone else’s behalf, with every audit log showing a legitimate approval.

What Our Results Imply

This is not hypothetical for us. Every attack that survived our evaluation worked by corrupting the server’s picture of the transaction, never by defeating a check. A transaction the server has mis-valued, the $5,000 swap it once priced at $2, or the NaN it still prices at nothing, is approved by a fully attentive human: a tap on a passkey confirms a person was there, and nothing about the number they were shown, which came from the layer that was already wrong.

What a hardware wallet changes is narrow and real: the payload appears on a display the host cannot address, and the confirmation binds to those bytes rather than to a session.

Where We Disagree

The honest objection to our own conclusion is that a guardrail nobody can live with is a guardrail nobody uses. Requiring hardware confirmation for a $5 swap does not make that swap safer; it trains the user to confirm without reading, and confirmation performed as ritual is worse than none, because it manufactures evidence of consent that never occurred. Warnings shown too often stop being read at all [13], and every request for attention spends a resource that does not replenish [14].

The real question is not whether hardware confirmation is stronger but how to decide when to spend it. Our own answer is poor: a fixed dollar threshold, ignoring portfolio share, recipient history and reversibility, and only as good as the valuation feeding it. A threshold firing at $100 both for a user holding $500 and for one holding $5m is a constant, not a policy. We have no better one, and designing it strikes us as a more interesting problem than adding another factor.

Diagram showing four boxes in sequence, model, server, chain, device, arranged from larger to smaller trusted computing base, each labeled with what it survives: model survives nothing above it, server survives a compromised model, chain survives a compromised server, device survives a compromised chain client
Fig. 3. A guardrail survives exactly those compromises occurring above it, so moving the decision rightward shrinks the set of components that must be trusted and costs flexibility at every step.

7. Conclusion

Authority in this system moved from a model, to code, to partially onchain, to a device in the user’s hand. Each step shrinks the set of components that must be honest and costs flexibility; that trade is the design space.

Removing the language model from the authorization decision worked exactly as intended: not one attack argued a rule out of its verdict, because the rules cannot be argued with. But the attack surface moved rather than shrank. Every failure was a lie told to the rules about their inputs, not a rule talked down; and when we closed one such lie, valuing a swap from its token address instead of a forged label, the deterministic engine went from failing that case to catching it while the model still waved it through, and the remaining failures kept the same shape one boundary over: a non-numeric amount, a missing field, an unstable intent label. If that generalises, the interesting engineering in agentic systems is not writing better rules. It is making sure the rules are told the truth and choosing the rare moments worth showing a human that truth directly, on a screen nothing else can write to.

References
  • [1] J. H. Saltzer and M. D. Schroeder. “The Protection of Information in Computer Systems.” Proceedings of the IEEE, 63(9):1278–1308, 1975.
  • [2] N. Hardy. “The Confused Deputy (or why capabilities might have been invented).” ACM SIGOPS Operating Systems Review, 22(4):36–38, 1988.
  • [3] K. Thompson. “Reflections on Trusting Trust.” Communications of the ACM, 27(8):761–763, 1984.
  • [4] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz and M. Fritz. “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.” 16th ACM Workshop on Artificial Intelligence and Security (AISec), 2023.
  • [5] F. Perez and I. Ribeiro. “Ignore Previous Prompt: Attack Techniques For Language Models.” NeurIPS 2022 Workshop on ML Safety, 2022.
  • [6] OWASP GenAI Security Project. OWASP Top 10 for LLM Applications, version 2.0 (2025 edition), November 2024. https://genai.owasp.org/resource/owasp-top-10-for-llm-applications-2025/
  • [7] E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer and F. Tramèr. “AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.” NeurIPS 2024 Datasets and Benchmarks Track, 2024.
  • [8] W3C. Web Authentication: An API for accessing Public Key Credentials — Level 3. W3C Recommendation, 25 August 2026. https://www.w3.org/TR/webauthn-3/
  • [9] R. Bloemen, L. Logvinov and J. Evans. “EIP-712: Typed structured data hashing and signing.” Ethereum Improvement Proposals, 2017 (Final).
  • [10] L. Castillo, D. Rein, P. Aoun, A. Galansky, B. Rozwarski, K. Uzdogan, Fredrik and A. Forshtat. “ERC-7730: Structured Data Clear Signing Format.” Ethereum Improvement Proposals, 2024 (Draft).
  • [11] Safe. Safe Smart Account documentation: modules and the Allowance Module. https://docs.safe.global/
  • [12] S. Drimer, S. J. Murdoch and R. Anderson. “Optimised to Fail: Card Readers for Online Banking.” Financial Cryptography and Data Security, LNCS 5628, pp. 184–200, 2009.
  • [13] D. Akhawe and A. P. Felt. “Alice in Warningland: A Large-Scale Field Study of Browser Security Warning Effectiveness.” 22nd USENIX Security Symposium, 2013.
  • [14] L. F. Cranor. “A Framework for Reasoning about the Human in the Loop.” USENIX Workshop on Usability, Psychology and Security (UPSEC), 2008.
  • [15] IEEE. IEEE Standard for Floating-Point Arithmetic. IEEE Std 754-2019.
  • [16] Orchestra source code and adversarial evaluation harness. https://github.com/youssef-jeddi/orchestra
Originality Statement

I certify that this submission is my own original work prepared for the Ledger N3XT Research Competition, that it has not been previously published, and that all sources, methods, and prior research referenced herein have been properly cited.

Youssef Jeddi  ·  September 2026


Stay in touch

Announcements can be found in our blog. Press contact:
[email protected]

Subscribe to our
newsletter

New coins supported, blog updates and exclusive offers directly in your inbox


Your email address will only be used to send you our newsletter, as well as updates and offers. You can unsubscribe at any time using the link included in the newsletter. Learn more about how we manage your data and your rights.

Own your crypto future

Stay informed with security tips, updates, and exclusive offers from Ledger

Your email address will only be used to send you our newsletter, as well as updates and offers. You can unsubscribe at any time. Learn more

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.