Signed, Not Sound
Research CompetitionWhy Hardware-Backed Custody Does Not Complete Agentic Authorization
AI agents are beginning to make payments without a person reviewing every transaction. Hardware-backed custody helps protect this process by keeping signing keys outside the agent’s runtime and, in some implementations, providing evidence about the device and software that produced a request. However, protecting the key does not prove that the agent understood what the user wanted. An agent can produce a valid, policy-compliant transaction that is still wrong for the user’s actual purpose. Protocols such as Google’s Agent Payments Protocol address part of this problem by checking cryptographically verifiable payment constraints before authorization, but errors can still occur when an agent interprets a task, selects a mandate or constructs an otherwise permitted action. This paper also distinguishes these authorization risks from device-level limitations. Sustained thermal pressure can reduce the throughput of mobile inference workloads, while low-bit quantization can reduce model accuracy. Whether either condition increases errors in payment-authorization payloads remains an open empirical question. To address these risks, the paper proposes layered architecture combining hardware-backed custody, independent deterministic authorization and optional evidence about the agent’s execution environment. It also outlines an experiment that measures availability, schema conformance and semantic correctness separately. The central conclusion is that hardware can establish which protected signer approved an action and, where attestation is available, provide evidence about measured execution conditions. Deterministic policy can determine whether the action satisfies explicit constraints, but neither mechanism alone proves that it reflects the user’s intent.
1. Introduction: Signed Does Not Mean Sound
When an AI agent can initiate a payment without human review, a compromised model, tool or intermediary can turn a manipulated instruction into a signed financial action. Research on third-party LLM routers demonstrates the importance of protecting the signing path. Liu et al. examined 428 third-party routers and found that nine injected malicious code into tool-call responses, seventeen accessed researcher-controlled canary credentials and one drained ETH from a researcher-controlled private key (Liu et al., 2026). Each router terminates its inbound TLS session, giving it plaintext access to prompts, credentials and tool-call payloads. Keeping the signing key outside this path reduces the risk of direct key theft, but a malicious intermediary may still alter a request before it reaches a legitimate signer.
Hardware-backed custody provides this isolation by placing cryptographic authority beyond the direct control of the model and its routing infrastructure. Ledger’s position that durable digital trust requires a physical root of trust captures the importance of establishing key control and trustworthy provenance (Gauthier, 2026). Yet custody is only the beginning of an authorization architecture. Once the key is protected, the system must still decide which proposed actions it should sign. This paper examines that second problem. It distinguishes custody, deterministic policy enforcement and evidence about the agent’s execution environment as separate layers of assurance. It then identifies a remaining semantic gap between satisfying formal constraints and fulfilling the user’s objective. The proposed architecture addresses this gap by combining hardware-backed signing with an independent, fail-closed authorization gate and escalation mechanisms for actions that cannot be approved confidently.
2. Background: Payment Protocols and Custody
Agentic payments bring together two related security problems. The first is custody: protecting the cryptographic keys used to approve and sign transactions. The second is authorization: deciding whether a proposed transaction is permitted under the user’s instructions. Emerging protocols address different parts of this process. Examining x402 and AP2 helps clarify what protocol-level checks can establish and where human intent remains difficult to represent.
2.1 Programmatic payment authorization
x402 and AP2 operate at different layers of the emerging agentic-payment stack. x402 provides payment messages and settlement schemes that allow a service to request, verify and settle a payment. Under its Exact EVM scheme, a facilitator checks the payer’s EIP-712 signature, EIP-3009 authorization parameters, recipient, amount, validity period, nonce, balance and transaction feasibility before settlement (x402 Foundation, 2025). Its verification boundary is the validity and settleability of the submitted payment payload.
AP2 focuses on delegated commercial authority. Google contributed AP2 to the FIDO Alliance in April 2026 as part of its community-led work on agentic authentication and payments, alongside the release of AP2 version 0.2 (FIDO Alliance, 2026). The specification uses SD-JWTs to secure Checkout Mandates and Payment Mandates. In autonomous transactions, a user approves an open mandate that defines the agent’s authority and applicable constraints. The agent later binds that authority to closed mandates representing a specific purchase. Verifiers then use deterministic code to validate mandate integrity and enforce the recorded conditions (Google Agentic Commerce, 2026).
Together, these protocols illustrate two forms of machine-verifiable control: x402 validates payment execution requirements, while AP2 represents and enforces delegated purchasing authority. Table 1 compares their verification boundaries with those of hardware custody and execution-context evidence.
| System | Primary control | What is verified | Residual limitation |
|---|---|---|---|
| x402 v2 | Scheme-specific signed payment payload and facilitator verification | For the exact EVM scheme: signer, recipient, amount, validity window, nonce, balance, and transaction feasibility | Verifies a valid payment authorization, not whether buying the resource was semantically appropriate for the principal |
| Google AP2 | SD-JWT-secured open or closed Checkout and Payment Mandates with deterministic constraint evaluation | The closed action is cryptographically bound to, and satisfies, the applicable user-approved mandate | Task interpretation and mandate selection remain outside the protocol; a permitted choice may still misunderstand the user’s objective |
| Hardware-backed agent custody | TEE, secure element, HSM, or threshold signing architecture | Depending on implementation: key control, code measurement, signing policy, or multi-party approval | Key or code provenance does not establish semantic correctness of the proposed action |
| Hardware-attested execution evidence | TEE or trusted system software producing signed measurements | Depending on implementation: measured code, device identity, model configuration and selected runtime conditions | Measurements require empirical calibration and do not independently establish semantic correctness |
2.2 Custody as a root of trust
Agent custody architectures protect signing keys from extraction and unauthorized use. Rather than storing a key inside the model runtime, these systems may place it in a trusted execution environment (TEE), hardware security module, secure element or threshold multi-party computation system (Alqithami, 2026). Each approach establishes a protected boundary between the software proposing a transaction and the mechanism authorized to sign it.
Some architectures distribute signing authority across multiple enclaves or participants. This arrangement prevents a single compromised environment from producing a valid signature independently (Acharya, 2025). One proposed local-agent identity architecture stores a device-bound key pair in a Secure Enclave or TPM and issues short-lived credentials without exposing the long-lived private key to the agent process (1Password, 2026).
These architectures differ in trust assumptions and operational control. TEEs and secure elements rely on hardware isolation, hardware security modules centralize protected key operations, and threshold systems divide authority among multiple participants. Despite these differences, each design separates probabilistic model execution from cryptographic signing authority.
3. The Residual Gap: Policy-Compliant but Semantically Wrong
Formal authorization policies define what an agent is permitted to do. A user’s objective defines what the agent is expected to accomplish. Although these may overlap, satisfying one does not necessarily satisfy the other.
For instance, an agent instructed to buy a remote-work laptop for less than $1,000 might be bound by a mandate limiting the retailer, product category, currency and total price. The agent could satisfy every condition while selecting a laptop with an incompatible operating system, insufficient memory or a delivery date later than the user’s deadline. It might also repeat a purchase already completed by another agent. Each transaction would comply with the recorded constraints without accomplishing the intended task.
This difference between permission and purpose creates the residual semantic gap. Constraints can exclude clearly unauthorized actions, but they cannot always represent compatibility requirements, contextual information, unstated preferences or the trade-offs a user would make when circumstances change. Adding more constraints can reduce ambiguity, but an open-ended mandate cannot anticipate every relevant condition. An authorization system must therefore distinguish among prohibited actions, clearly acceptable actions and uncertain actions that require further evidence or human confirmation.
Related problems appear in multi-agent security research. Tallam describes authorization-propagation failures in which delegated permissions, aggregated information or validity periods drift as a task moves through a workflow, even when no component is obviously compromised (Tallam, 2026). Uchibeke proposes a deterministic pre-action authorization system that places an independent policy gate between an agent and its tools. The gate evaluates each proposed tool call and issues a signed decision before execution (Uchibeke, 2026). This prevents the reasoning model from defining and exercising its own authority without independent review.
Hardware-backed guardrail research establishes a similar boundary. Proof-of-Guardrail can attest that specified guardrail code executed inside a TEE, but its authors distinguish evidence of execution from evidence that the guardrail was effective in a particular case (Jin et al., 2026). Hu and Le likewise show that hardware isolation and application-level guardrails address different failure modes and should be combined (Hu & Le, 2026).
The distinction can therefore be stated precisely: provenance identifies the origin and execution conditions of an action, deterministic policy evaluates compliance with explicit rules, and semantic evaluation considers whether the action advances the user’s objective. Section 5 assigns these responsibilities to separate architectural controls.
4. What On-Device Inference Research Actually Establishes
Sustained inference on passively cooled mobile systems imposes a measurable performance cost. Tummalapalli et al. benchmarked a 4-bit Qwen 2.5 1.5B model over twenty repeated iterations on an iPhone 16 Pro and Samsung Galaxy S24 Ultra. Under the tested hardware and software configurations, the iPhone’s sustained throughput plateau was 41.5% below its peak, while the S24 plateau was 15.0% below its peak (Tummalapalli et al., 2026). These findings establish platform-specific throughput degradation under sustained load. They do not establish corrupted arithmetic, declining semantic accuracy or workload termination as a general thermal effect.
Quantization is a separate concern. It is selected when a model is prepared or loaded; it is not caused by thermal throttling. Chen et al. cite Qualcomm results showing substantial accuracy degradation for some small edge models under INT4 and INT8 configurations, while SPEED-Q demonstrates that carefully designed low-bit methods can materially outperform earlier 2-bit approaches across vision-language benchmarks (Chen et al., 2025; Guo et al., 2026). These studies establish that model size, precision, method, and task can affect accuracy. They do not measure payment authorization or structured tool calls.
The open research question is therefore deployment-level rather than a claim that heat directly makes digital computation incorrect. Sustained load may increase latency, timeouts, truncation, fallback behavior, memory pressure, or OS termination. A lower-bit model may differ in schema or semantic accuracy from a higher-precision baseline. Their interaction may matter when a system changes models, token budgets, or execution paths under resource pressure. No cited study establishes that thermal state alone degrades authorization reasoning. The contribution here is to identify and make that hypothesis falsifiable.
5. A Layered Authorization Architecture
A complete authorization architecture should contain three distinct layers. The first is custody assurance, which protects the signing key and may attest the signer, device and measured code. The second is an independent deterministic authorization service, which evaluates every proposed action against the principal’s delegated constraints. This service should remain outside the model’s control, fail closed when required policy or evidence is unavailable, and return a signed allow, deny or escalate receipt. The third layer is optional execution-context evidence, which can support auditing and escalation once its relationship to observed failures has been empirically calibrated.
Execution-context evidence should not be presented as proof of reasoning integrity. Hardware or trusted system software may report a model identifier or hash, quantization configuration, thermal state, memory pressure, structured-output validation result, policy version and action-receipt hash. Implementations must clearly distinguish hardware-attested measurements from metadata reported by ordinary software. They should also disclose as little information as necessary because model identifiers, device conditions and policy metadata may enable fingerprinting or reveal sensitive operational details.
Authorization decisions should not depend on an opaque reliability score. Financial systems require predictable and auditable outcomes. Once testing establishes meaningful thresholds, deterministic policy can map specific conditions to specific actions. Malformed output may be denied; a timeout may trigger a retry after cooldown; disagreement between semantic checks may route the request to a stronger model; and a high-value payment to a new recipient may require confirmation on a trusted display. Until such thresholds are calibrated, thermal state should be recorded as research evidence rather than interpreted as proof that a decision is safe or unsafe.
6. Methodology and Proposed Evaluation
This section separates the evidence supporting the paper’s current argument from the empirical work needed to investigate its open questions. Section 6.1 explains how existing sources and findings are used, while Section 6.2 presents a controlled benchmark for measuring the effects of quantization and sustained load on structured authorization.
6.1 Methodology of this paper
This paper presents a position supported by a synthesis of existing research rather than new experimental results. It examines primary protocol specifications alongside studies of router security, deterministic authorization, trusted execution, sustained mobile inference and low-bit quantization. The evidence is classified as established protocol behavior, measured security or performance findings, and hypotheses that require further testing. In particular, the paper does not treat a relationship between thermal state and structured-output correctness as an observed result.
6.2 Proposed evaluation
The proposed benchmark uses a factorial design that varies two conditions independently: model quantization and sustained computational load. Base model version, prompts, tool schema, decoding parameters, token budget and task distribution remain fixed. Each task requires the model to produce a JSON payment-authorization payload for which the correct recipient, amount, asset, constraints and authorization decision are known in advance. Trials should be conducted under both a cooled baseline and sustained load, using randomized task order, repeated sessions and multiple devices. Cooldown periods between selected blocks allow the system to return toward its baseline condition. Randomization and cooldown are important because thermal state may otherwise be confused with elapsed runtime, battery condition or task order.
Runtime measurements should include thermal state, memory pressure, latency, timeouts, truncation, retries, fallback behavior, operating-system termination and energy consumption where the device exposes it. Evaluation outcomes should be reported separately in four categories: runtime availability, schema conformance, semantic correctness against the ground-truth oracle and acceptance by an independent deterministic policy engine. This separation prevents successful execution, a well-formed payload or policy acceptance from being counted automatically as semantic correctness. The analysis should test whether sustained load is associated with availability failures, whether quantization changes schema or semantic accuracy, and whether the two factors interact. A mixed-effects model can account for differences among devices, sessions and tasks. Results should include confidence intervals and effect sizes rather than relying only on point estimates. Before testing, the study should define the smallest change in correctness that would matter operationally and use enough trials to detect it reliably. If the resulting confidence interval excludes changes of that size, the study may conclude that no practically meaningful effect was found under the tested conditions. A nonsignificant result by itself should be treated as inconclusive.
Hardware Roots of Trust and the Authorization Gate
Ledger’s position that digital trust must ultimately rest on a physical root rather than remotely mutable software is directly relevant to agentic commerce (Gauthier, 2026). Router attacks demonstrate the practical value of removing signing keys from the model’s software and network path. When a private key is generated, stored and used exclusively inside a protected hardware boundary, routers and model runtimes do not receive the raw key material needed to sign independently.
Extending this architecture to autonomous agents requires pairing the protected signer with a narrow deterministic policy engine. Hardware protects cryptographic authority, policy restricts the permitted execution scope, and signed receipts create an auditable record of each decision. Trusted displays or equivalent confirmation mechanisms can then return control to the principal when a transaction exceeds established limits or introduces an unfamiliar recipient.
This division of responsibility follows Ledger’s architectural principles without asking secure hardware to evaluate open-ended reasoning. The model proposes an action, deterministic policy allows, denies or escalates it, and hardware protects the keys while attesting selected execution facts. The principal retains explicit control over spending limits, recipients and escalation conditions. This division applies Ledger’s hardware-rooted trust model while preserving an independent authorization boundary.
7. Conclusion
Hardware-backed custody provides a necessary foundation when agent runtimes, routers or service operators might expose signing authority. Current protocols such as AP2 demonstrate that signatures should be combined with deterministic verification of delegated mandates and transaction constraints.
Mobile-inference research adds a concrete systems question to this architecture. Sustained thermal load reduces throughput under the tested mobile configurations, while quantization can affect task accuracy. Their effects on structured authorization payloads have not yet been measured. A rigorous benchmark should therefore evaluate availability, schema conformance and semantic correctness as separate outcomes.
Until those results exist, the defensible architecture is layered: hardware-backed custody, independent deterministic authorization, optional calibrated execution evidence and fail-closed escalation. Signed is necessary. Sound must be demonstrated through policy, testing and accountable system design.
- 1. Gauthier, P. (2026). Revenge of the Atoms. Ledger. https://www.ledger.com/blog-revenge-atoms
- 2. Liu, H., Shou, C., Wen, H., Chen, Y., Fang, R. J., & Feng, Y. (2026). Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain. arXiv:2604.08407. https://doi.org/10.48550/arXiv.2604.08407
- 3. x402 Foundation. (2025). x402 Protocol Specification, Version 2. Retrieved August 31, 2026. https://github.com/x402-foundation/x402/blob/main/specs/x402-specification-v2.md
- 4. Google Agentic Commerce. (2026). Agentic Payment Protocol (AP2), Version 0.2: Protocol Specification and Agent Authorization Framework. Retrieved August 31, 2026. https://github.com/google-agentic-commerce/AP2/blob/main/docs/ap2/specification.md
- 5. Tallam, K. (2026). Authorization Propagation in Multi-Agent AI Systems: Identity Governance as Infrastructure. arXiv:2605.05440. https://doi.org/10.48550/arXiv.2605.05440
- 6. Uchibeke, U. (2026). Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents. arXiv:2603.20953. https://doi.org/10.48550/arXiv.2603.20953
- 7. Alqithami, S. (2026). Autonomous Agents on Blockchains: Standards, Execution Models, and Trust Boundaries. arXiv:2601.04583. https://doi.org/10.48550/arXiv.2601.04583
- 8. Acharya, V. (2025). Secure Autonomous Agent Payments: Verifying Authenticity and Intent in a Trustless Environment. arXiv:2511.15712. https://doi.org/10.48550/arXiv.2511.15712
- 9. 1Password. (2026). Delegated Authority, Running Locally: Give an Agent on Your Machine an Identity You Can Trust. https://1password.com/blog/ai-agent-identity-delegated-local
- 10. Hu, B. S., & Le, Q. D. (2026). AI Agents Need Both Hardware-Backed Security and Application-Level Guardrails. Proceedings of the 9th Workshop on System Software for Trusted Execution (SysTEX ’26), 92-95.
- 11. Jin, X., Duan, M., Lin, Q., Chan, A., Chen, Z., Du, J., & Ren, X. (2026). Proof-of-Guardrail in AI Agents and What (Not) to Trust from It. arXiv:2603.05786. https://doi.org/10.48550/arXiv.2603.05786
- 12. Chen, L., Feng, D., Feng, E., Wang, Y., Zhao, R., Xia, Y., Xu, P., & Chen, H. (2025). Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference. Proceedings of SOSP ’25, 359-374. https://doi.org/10.1145/3731569.3764808
- 13. Tummalapalli, P., Arayakandy, S., Pal, R., & Kundan, K. (2026). LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load. arXiv:2603.23640. https://doi.org/10.48550/arXiv.2603.23640
- 14. Guo, T., Zhao, S., Zhu, S., & Ma, C. (2026). SPEED-Q: Staged Processing with Enhanced Distillation Towards Efficient Low-Bit On-Device VLM Quantization. Proceedings of AAAI 2026. arXiv:2511.08914. https://doi.org/10.48550/arXiv.2511.08914
- 15. FIDO Alliance. (2026, April 28). FIDO Alliance to Develop Standards for Trusted AI Agent Interactions. https://fidoalliance.org/fido-alliance-to-develop-standards-for-trusted-ai-agent-interactions/
I confirm that this paper is my own original work. AI tools were used for literature discovery and editorial assistance. I independently verified the cited sources and take responsibility for the paper’s analysis, argument, and final text.
Huzaifa Bin Hamid · 31 August 2026