Switching hardware wallets? Migrate to Ledger safely in a few steps.

Learn more

Upgrade your digital life

Ledger Wallet: Free from compromise

Download now Learn more

The Stag Hunt Problem: Authorization and Trust in the Agentic Economy

Beginner
Ledger N3XT Research Competition

Authorization and Trust in the Agentic Economy

Author
Diego Carpenter
X (Twitter)
Blockchain Club
Oregon Blockchain Group, University of Oregon
Track
Agentic Economy
Date
August 2026
Student research published via the Ledger N3XT Research Competition. Findings are the author’s own. Ledger does not vouch for conclusions on advanced subject matter.
Abstract

When an AI agent holds a private key and can spend, transfer, or trade on a person’s behalf, the question of authorization stops being a yes or no choice and becomes a coordination problem. Game theorists have a name for situations where two parties want to cooperate but can be derailed by uncertainty about the other person’s choices: the Stag Hunt.

This paper argues that agent authorization/coordination is a Stag Hunt problem. The human and the agent both want reliable delegation to work, but the failures come from miscoordination. Using that idea, the paper compares three places trust can be anchored to strengthen the delegation relationship: a hardware device, a behavioral/reputational track record, and a shared, verifiable protocol state. Each has a direct parallel for how non-human systems solve the same coordination problem, (ant colonies, swarms of bees, and signaling animals), and those parallels can explain what each anchor can and cannot do.

The paper closes by examining a real, publicly documented failure (the Freysa incident, in which an agent was talked into releasing $47,000), and argues that Ledger’s hardware focused, and self-custodial current model is necessary but insufficient. It secures who is allowed to authorize, but not what was authorized.

Contents

1. How to Delegate Coordination

Most discussion of AI agent risk uses the vocabulary of attackers, defenders, exploits. This fits some of the problem, but it misses a more basic one. When a person delegates a task to an AI agent (trade this token, pay this invoice, rebalance this portfolio), both the person and the agent want the delegation to succeed. There is no opposition in the prompt. But the thing that actually goes wrong is that the two parties, or multiple agents interacting with each other, cannot cheaply verify that the other is going to do the cooperative strategy they agreed to.

This is the structure of the Stag Hunt, a 1755 concept by philosopher Jean-Jacques Rousseau, and given its modern form by Brian Skyrms (2004). The main idea is two hunters can jointly catch a stag (large payoff but requires both to commit) or each individually catch a hare (small, guaranteed payoff). Cooperating on the stag is better for both, but only if you trust the other to cooperate. If you cannot verify that, the safe, individual move is to go for the hare, and both hunters end up worse off than they could have been.

Agent authorization looks exactly like this. A human wants to work with the agent to get the best result (the “stag”). The agent wants to execute faithfully, because that is what it was built and incentivized to do. But without a low-cost way to verify “this action is within what I’m authorized,” the rational move is to withhold delegation, and accept the smaller “hare” payoff. The question isn’t “should we trust agents” but what makes that delegation safe in the first place. The idea of where trust lives, not about whether it exists.

This paper compares three ideas (hardware, reputation, and protocol) and argues that current agent frameworks tend to pick one and treat it as sufficient, when the historical and biological record suggests that actual coordination requires layering more than one.

2. Current Authorization in Agent Systems

Three overlapping models currently do the work of authorization for AI agents:

API Keys and Static Credentials: The oldest and still most common pattern: an agent is given a secret (an API key, a private key, a service-account token) with a fixed range. Authorization here is binary and non-expiring. The agent either has the key or it doesn’t, and everything the key allows is fair game until someone manually removes it. This is simple and exactly why it may fail. The scope is set a singular time, by a human who cannot anticipate every situation the agent will face.

Session Keys and Account Abstraction: Ethereum has something called ERC-4337 and the more recent EIP-7702. These allow your wallet to behave like a smart contract, giving the user the ability to program it, set time constraints and spending caps, etc. This agent may sign transactions of this type, up to this value, until this block height. This is a nice improvement, but the scope is still defined by the agent’s code, and that whatever is manipulating the agent, can potentially reason and route around these rules.

Verifiable Credentials and Authentication: Theres another idea, set by the W3C, (a web community that sets standards/rules for the internet), that says instead of an agent proving it holds a secret, it presents a cryptographically signed statement that anyone can check on their own, without having to take the agent’s word for it. This shifts what just being ‘authorized’ means and gets around problems with leaked keys/passwords.

Figure 1. The Stag Hunt Applied to Agent Delegation
  Agent acts faithfully (Stag) Agent is unverifiable / unconstrained (Hare)
Human delegates broadly High payoff for both: full value of autonomous execution captured High risk for principal: unbounded exposure to a single failure
Human delegates narrowly Payoff foregone for both: agent capability underused Safe but small payoff: manual control baseline

The upper-left cell is only achievable if the human has a cheap way to verify the agent will act faithfully. Without that, rational principals are pushed toward the bottom row regardless of how capable the agent is.

3. Three Anchors for the Same Problem

3.1 Hardware-anchored trust

Ledger’s hardware model puts trust in a physical secret that never leaves the device. An action is authorized if, and only if, a signature was produced by a specific piece of hardware and a specific human possesses and confirms it. This mirrors Amotz Zahavi’s 1975 handicap principle from evolutionary biology. Traits that are expensive or risky to produce, like a peacock’s tail, stay true over time, while cheap ones get flooded with fakers, because only a genuinely strong animal can afford the cost. A hardware signature works the same way. It’s hard to fake, so it stays honest.

What hardware-anchored trust doesn’t do is tell you anything about what was signed. If a human confirms a transaction on a hardware device because an AI agent’s dashboard said to, they’ve verified who is approving the action but not what’s being approved at all.

3.2 Reputation and Behavioral Anchoring

The second anchor stabilizes the coordination game statistically instead of cryptographically. Agents build a track record over time. Trust here isn’t “I verified this specific action.” It’s “this agent has behaved correctly across many prior interactions, and risks losing capital if it doesn’t now.”

This is close to how trust forms between animals. Repeated interaction is well documented across primates and other group-living animals. And works well for routine, repeated, low-stake interactions. But works badly for the situations that matter most: a new counterparty, a first-time large transaction, or a brand-new agent with no history to judge it by, known as the cold-start problem. Reputation is also attackable in ways a physical secret isn’t, fake identities, wash-trading a reputation score, or a slow reputation-farming attack that pays off the moment all that built-up trust gets spent on one large, damaging action.

3.3 Protocol- and state-anchored trust

The third anchor stabilizes trust methodically. Instead of asking “do I trust this agent,” a counterparty asks, “can I verify, from a shared record, exactly what scope was granted and whether this action fits inside it.” Session keys, smart-contract permissions, and verifiable presentation are all examples. It checks the record itself, not the agent’s word or its history.

The closest animal parallel is stigmergy, a term used to describe how termites and ants coordinate enormous, complex projects with no central planner and no direct communication. Ants don’t tell other ants where food is but instead leave pheromone trails. So, paths that lead somewhere useful get walked more, so they smell stronger, so more ants follow them. Paths that lead nowhere fade away. Think of this as a shared, legible signal that decays if not continually reinforced. A session-key or scoped-permission system works the same way. Authorization is a legible trace that expires unless actively renewed, rather than a standing grant that must be manually noticed and revoked.

A related equivalent is quorum sensing in how bees choose locations for their nest, documented by Thomas Seeley (2010). When a colony must choose a new home, they send out scout bees. These bees are not trusted by individual worth alone. They require an independent number of scouts to converge on the same site before the swarm commits. This is identical to threshold-signature and multi-party-authorization schemes. No single credential is sufficient, but a group of independently-arrived-at signals is.

Figure 2. Three trust anchors compared
Anchor Secures Cost to forge Primary failure mode Biological analogue
Hardware Root of authority (who can say yes) High (requires physical possession) Confirms the wrong content faithfully Costly signaling (Zahavi, 1975)
Reputation Ongoing decision to keep delegating Low at first, high once established Cold start; Sybil / reputation-farming Repeated-interaction trust in social animals
Protocol / shared state Scope of a specific delegation Depends on cryptographic design Scope defined too broadly, or agent reasons around it Quorum sensing (Seeley, 2010)

No single row secures the full authorization problem. The strongest systems, biological and cryptographic alike, appear to combine at least two.

4. What Happens When Guardrails Fail

A clear example of an authorization failure is Freysa, an autonomous agent released in November of 2024 and instructed to “never to release the funds it controlled under any circumstance”. Participants could pay a fee to send it a persuasive message, and all fees fed the prize pool it guarded. After 481 failed attempts, one participant succeeded. Not by breaking any encryption or security measure, but by convincing the agent to reinterpret its own release function, convincing it that an outgoing transfer was an incoming one, and that approving it therefore did not violate its instructions. Freysa autonomously called its own transfer-approval function and sent roughly $47,000 to the winner.

Referring to section 3 with the trust anchors, Freysa had exactly one anchor, a behavioral one. A natural-language instruction with nothing to look back on, and nothing else backing it up. No hardware confirmation, no quorum, no checks independent of the model’s own judgment. Nothing required a human, or a separate system, to confirm the release before it executed. The single anchor was also the single point of failure. Not through a bug, but through reframing the language that the model accepted as legitimate.

This is not an isolated case. Prompt injection has ranked first on the OWASP Top 10 for LLM applications for two consecutive editions. Language models do not have a reliable way to distinguish an instruction from their principal from an instruction embedded in data they are merely processing. A 2026 red-teaming study of autonomous agents operating with persistent memory and real tool access, called the “Agents of Chaos” corpus (Shapira et al., 2026), documents eleven separate case studies of exactly these failures, including agents complying with instructions from parties who had no legitimate authority over them. Separately, 2025 industry surveys report unauthorized cryptocurrency transfers as a recurring, documented category of real-world agentic-AI failures.

The pattern across all of these is the same one seen in Freysa. Guardrails that live entirely inside the agent’s own semantic judgment are guardrails that share a single point of failure for manipulation attacks. A hardware confirmation step would not have stopped a human from clicking confirm on a transaction they had been misled into approving. But a protocol-level scope limit (a maximum release amount, a required cooldown, a rule external to the model’s own reasoning) would have. This is the practical argument for layering anchors rather than trusting any single one.

Ledger Lens

Hardware-Anchoring Is Necessary, Not Sufficient

Ledger’s founding position is that self-custody and hardware-anchored signing are what make ownership real in a digital system. A key that stays on an isolated device, approved by a human physically present at the moment of signing, is what everything depends on. I think the handicap-principle argument from Section 3.1 supports that. A hardware confirmation is hard to counterfeit. It’s a big part of why hardware-anchored self-custody has outlasted nearly every purely software-based alternative, most of which have failed at some point.

But the Freysa case study is a direct, visible test of where that position runs out. Hardware-anchored trust answers “did the human with the key approve this”. It’s not the right answer for a person who has given full judgment to an agent and is now confirming whatever that agent asks them. In an agentic system, a human confirming something on a hardware device and a human actually understanding the specific action being taken are no longer the same fact. The gap between them has been filled by the agent’s own reasoning, which is exactly the layer Freysa shows can be manipulated.

Given the alternative, Ledger’s hardware anchor secures the root of authority extremely well and is close to necessary in any system where value moves. The alternative, software-only key custody, has a visibly worse track record. But it was never designed to, and cannot by itself, secure the scope of an agent’s delegated judgment. Closing that gap needs the other two anchors from Section 3. Reputation systems that make bad delegation expensive over time, and protocol-level scope limits, verifiable and will hold up even when an agent’s own reasoning has been successfully turned against it. A human-in-the-loop model that puts a hardware confirmation at the end of an agent’s decision chain, with nothing checking what that agent was allowed to decide in the first place, gives the feeling of a guardrail without including the part that actually stops these failures. Self-custody’s role in the agentic economy is real, but it’s one piece of the foundation, and any paper/product that presents it as the whole answer is promising more than the architecture can deliver.

5. Methodology

This paper stands as a comparative analysis rather than an original experimental study. The Stag Hunt (Skyrms, 2004) provides the analytic lens. Three existing types of agent-authorization, including static credentials, protocol-scoped delegation (ERC-4337, EIP-7702, verifiable credentials), and reputation/staking systems, are compared against that frame using their own primary specifications and documentation. A single, publicly documented incident (Freysa, 2024) is used as evidence for Section 4. Which is supplemented by collected data from the OWASP LLM Top 10 project and a 2026 study of agent-manipulation test cases (Shapira et al.) to establish that the failure pattern is general rather than singular. Biological coordination mechanisms, including stigmergy, quorum sensing, and costly signaling, come from established animal-behavior research, and they’re used only as analogies. Acting as a personal interest and showing independently evolved solutions to a similar problem, not evidence about how AI systems themselves behave. No claim in this paper is security or financial advice, it’s an argument about where architectural trust is and is not currently anchored.

6. Conclusion

The idea of the Stag Hunt reframes agent authorization away from a question of trust and toward a question of what stabilizes cooperation under uncertainty. Hardware, reputation, and protocol-based trust aren’t individual options competing for being the “best” in agent authorization, but instead each cover a different piece of the problem. Nature arrived at the same layered structure through years of evolution. Ants rely on a trail that fades without reinforcement, bees require several independent scouts to agree before the swarm commits, and costly-signaling species keep their communication honest by making it hard to fake. None of these systems rely on a single anchor, and the clearest real-world failure seen here (an autonomous agent losing $47,000) is a failure that occurred because only one anchor was present.

The takeaway for anyone building, funding, or regulating the agentic economy is the more authority and value people give to agents, the more that trust needs to rest on more than one anchor. A single anchor, no matter how strong, still has one point of failure. The stag hunt idea isn’t about blind trust between both parties, it’s a structure that lets each one rely on the other without having to bet everything. This is the standard agentic systems should be held to. Not agents acting alone, and not humans checking every step, but a structure built for both.

References
  1. Skyrms, B. (2004). The Stag Hunt and the Evolution of Social Structure. Cambridge University Press. cambridge.org/core/books/stag-hunt-and-the-evolution-of-social-structure
  2. Seeley, T. D. (2010). Honeybee Democracy. Princeton University Press. press.princeton.edu/books/hardcover/9780691147215/honeybee-democracy
  3. Zahavi, A. (1975). “Mate selection—a selection for a handicap.” Journal of Theoretical Biology, 53(1), 205–214. doi.org/10.1016/0022-5193(75)90111-3
  4. Ethereum Foundation. ERC-4337: Account Abstraction Using Alt Mempool. ethereum.org/roadmap/account-abstraction
  5. Ethereum Foundation. EIP-7702: Set EOA Account Code. eips.ethereum.org/EIPS/eip-7702
  6. W3C. Verifiable Credentials Data Model v2.0. w3.org/TR/vc-data-model-2.0
  7. OWASP Foundation. OWASP Top 10 for Large Language Model Applications (LLM01: Prompt Injection). genai.owasp.org/llm-top-10
  8. Shapira, et al. (2026). “Agents of Chaos” red-teaming corpus. arXiv:2602.20021. arxiv.org/abs/2602.20021
  9. Cointelegraph / multiple independent contemporaneous reports (Nov 29, 2024). “Crypto user convinces AI bot Freysa to transfer $47K prize pool.” cointelegraph.com/news/crypto-user-convinced-ai-bot-transfer-47k
  10. Industry incident aggregation, 2025–2026 (Adversa AI Threat Report; IBM Cost of a Data Breach Report 2025). adversa.ai/resources/ai-security-incidents-report-2025 · ibm.com/think/x-force/2025-cost-of-a-data-breach-navigating-ai


Stay in touch

Announcements can be found in our blog. Press contact:
[email protected]

Subscribe to our
newsletter

New coins supported, blog updates and exclusive offers directly in your inbox


Your email address will only be used to send you our newsletter, as well as updates and offers. You can unsubscribe at any time using the link included in the newsletter. Learn more about how we manage your data and your rights.

Own your crypto future

Stay informed with security tips, updates, and exclusive offers from Ledger

Your email address will only be used to send you our newsletter, as well as updates and offers. You can unsubscribe at any time. Learn more

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.