The Last-Mile Oracle – Autonomous Agents in the Physical World
Research CompetitionWhy Agentic Commerce Needs a Graduated, Stakes-Based Path to Verifying Physical-World Outcomes
AI agents now operate autonomously negotiating prices, authorizing payments, settling transactions, and much more. This works because software is verifiable; for example an agent can prove that an API returned the correct data or that a smart contract executed correctly. The physical world does not have this same system of proof. Suppose a “home” agent needs to fix a broken sink, It cannot verify that this problem exists and it cannot prove that this task is completed. This is the core limitation of agentic commerce. Current solutions address this poorly and are vulnerable to subjectivity and market manipulation (Eskandari et al., 2021). In this paper it will use Graduated Verification Framework in place of an all-or-nothing trust model by threat-modeling three scenarios, home repair, last-mile delivery, and physical data collection. Using this we can see how the amount of proof a transaction requires scales with its stakes and reversibility. Low-value, easily-reversed transactions need only a single automated signal. While high-value or irreversible ones require stronger evidence. We will be able to achieve from device telemetry and decentralized physical infrastructure networks (DePIN) (Sarkar, 2023) up to human confirmation as the final backstop when automated proof isn’t sufficient.
1. Introduction
You are a homeowner and your HVAC system is failing, now consider that before you even know this problem exists a home autonomous agent has sourced a local technician and escrowed the payment; Through the current identity the agent can verify the technicians digital credentials but the problem occurs in the verification of a job completed: the agent can not verify that this repair actually took place or was even attempted.
This gap has a name: the Last-Mile Oracle problem. We define Oracle as any mechanism that brings real world information on chain, readable and verifiable. The last-mile oracle problem is the specific case where that information is not something simple like a weather reading in a location but that is a confirmation that a physical task was completed in the real world. Every agentic transaction that reaches into the physical world follows the same loop: an agent identifies a need, sources a provider, authorizes payment, and receives verifiable outcome. Current infrastructure does the first 3 but cant do the 4th.
The current barrier to fully autonomous machine-to-human commerce. While smart contracts excel at managing deterministic digital state they are always executed in a closed environment. Thus fundamentally blind to off-chain reality (Eskandari et al., 2021). Agentic infrastructure today attempts to standardize digital identity and payment settlement but their is no comparable push for verification of physical outcomes.
This paper evaluates the current agentic commerce stack using a comparative threat-modeling approach against three physical-world scenarios. We show that relying on single-point attestation or subjective crowdsourced oracles is insufficient for autonomous commerce. Instead, we propose a layered architecture where the burden of proof scales with the financial and safety stakes of the transaction, drawing on verifiable hardware telemetry (Sarkar, 2023) and establishing human-in-the-loop physical confirmation as the necessary boundary for irreversible, high-stakes agentic actions.
2. The Current Stack, and Where It Stops
There are currently three protocols that form the backbone of agentic commerce, each solving an individual piece of the transaction lifecycle.
The first step is discovery: ERC-8004 gives agents a portable, on-chain handle. This makes them discoverable to other agents, where they can search, filter, and select a counterparty before initiating a transaction. It functions with the agent registering as an ERC-721 token whose metadata points to a registration file. The file describes its functions and capabilities, also including the reputation registry where counterparties can post feedback onchain. Another component is the validation registry which allows an agent’s work to be checked through stake-secured re-execution, zero-knowledge machine learning proofs, or trusted-execution-environment attestation. These mechanisms basically just confirm that a computation ran correctly (software wise). Nothing about physical task completion, thus no re-execution path exists for something like a plumbing repair. The standard’s own authors even acknowledge this. The state that registration cryptographically ties an identity to its metadata cannot guarantee the agent’s advertised capabilities are functional or that its claimed work was real.
The second step is authorization: Google’s Agent Payments Protocol (AP2) and Visa’s Trusted Agent Protocol (TAP) focus on proving agent actions within human decision making. AP2 does this through a chain of cryptographically signed Mandates. These mandates include Intent, Cart, and Payment. An Intent Mandate captures the user’s original instruction and its constraints this would be things like a price cap, approved merchant list, etc. A Cart Mandate will lock in the specific items and price. A Payment Mandate then authorizes the transfer itself, carrying a hash of the matched Intent and Cart so a card network or stablecoin processor can verify the chain without seeing the full transaction detail. This solves the problem with unauthorized spend in the agentic company and will guard against some agent hallucination/rouge agents. AP2 lets us prove a user agreed to pay for a sink repair. It cannot “prove” the sink was repaired.
The third step is settlement: this is what x402 is for. x402 handles movement of stablecoin value by letting an agent complete a machine-to-machine payment directly inside web and API requests. This system is in production builds today being used by many. In the last 30 days (08/14/2026) it has processed ~75M transactions and $24M USD in volume. This shows that this service is very main stream and will continue to grow over time. The only problem with this system is settlement speed as it goes against what we want. x402 is built to release funds quickly and automatically, which is the wrong default when the thing being paid for cannot yet be verified. Wang (2026) catalogs several classes of exploits specific to this speed, like attacks at the discovery and payment-presentation stages that succeed precisely because settlement completes before any outcome check occurs.
These three layers answer: Who is this agent, did a human authorize this spend, and did the money move. All of these miss: Did the paid-for outcome actually occur. This gap becoming blatant once an agent’s counterparty is a human performing physical work in the real world.
The industry is not entirely blind to this gap. Goenka et al. (2026) propose TessPay, a “verify-then-pay” architecture that locks funds in escrow during task execution and releases them only once evidence satisfies a verification predicate. This is a direct critique of the settlement-first default this section has described. But TessPay is still about software only; they can attest that an API call returned a specific response or that a computation ran inside a secure enclave. But this system can not attest that a technician tightened a valve. TessPay’s existence confirms that settlement-before-verification is a real design flaw and also proves how we have systems in place to do all but the last verification layer we mention here.
3. Threat Model Scenarios
To map the limits of agentic commerce we have to look at some examples of agents utilizing physical labor. We will evaluate our threat model against three specific scenarios:
- Home Repair (Variable Judgment): An agent detects a failing appliance. It dispatches a human technician to perform a physical repair inside a residence.
- Last-Mile Delivery (Spatial Verification): An agent purchases a physical good. It requires proof that it was deposited at a specific geographic coordinate.
- Physical Data Collection (Sensor Attestation): An agent pays a human to photograph, measure, or audit a physical piece of infrastructure (e.g., verifying a billboard was placed).
These scenarios allow us to evaluate where current verification mechanisms succeed and where they fail when confronting the variables of the physical world.
4. Threat-Modeling the Verification Mechanisms
The fundamental vulnerability in agentic commerce lies at the Operational Technology (OT) and Information Technology (IT) boundary. Smart contracts and AI agents exist entirely within the IT layer as in they process deterministic digital state. Physical actions occur in the OT layer; they are subject to physics, human judgment, and environmental friction. Moving this data reliably from the OT layer to the IT layer is the essence of the Oracle Problem (Eskandari et al., 2021). The current solutions in agentic commerce fall into three categories of attestation.
Media Attestation (Subjective Visual Proof)
The simplest verification method is requiring a human worker to upload a photo of the completed task (e.g., a repaired sink or a delivered package). However, digital media can be easily spoofed (images can be staged, reused from prior jobs, or synthetically generated, etc.) especially today without hardware-level cryptographic binding. While device-level attestation (ensuring the photo was taken on a specific device at a specific time) would raise the cost of an attack, it only proves a photo was taken. So this solution is simple in practice and could work in some situations it is easily schemeable.
Subjective Oracles and Escrow (Game-Theoretic Proof)
A prevailing Web3 approach to dispute resolution is the subjective oracle, utilized by platforms like Kleros or UMA. In this model funds would be held in escrow and if the agent were to dispute an outcome, a decentralized pool of human jurors votes on the truth. While mathematically elegant, crowdsourcing subjective truth merely shifts the trust problem from the worker to the jurors. More critically, it violates the latency requirements of closed-loop industrial and agentic systems. Autonomous operations require deterministic execution; relying on a multi-day human voting process delays finality and makes autonomous scaling impossible.
Objective Hardware and DePIN (Deterministic Proof)
Decentralized Physical Infrastructure Networks (DePIN), such as Hivemapper or DIMO bridge the physical gap by relying on hardware rather than human judgment (Lin et al., 2024). A vehicle’s dashcam deterministically proves a road was mapped. However, DePIN’s problem space is narrower than generalized agentic commerce. DePIN relies on fixed sensors performing predictable tasks. Conversely, a human repairing a sink exercises variable judgment that a simple GPS or accelerometer ping cannot independently verify.
| Verification Mechanism | Home Repair | Last-Mile Delivery | Data Collection | Core Failure Mode | Confidence |
|---|---|---|---|---|---|
| Media Attestation | Low (quality unseen) | Medium (shows package) | Low (easily spoofed) | Deepfakes, staged photos, reused assets | Low |
| Subjective Oracles (Escrow) | Medium (jurors vote) | Medium (jurors vote) | Medium (jurors vote) | High latency, shifts trust to third-party voters | Medium |
| Objective Hardware (DePIN) | Low (cannot verify skill) | High (GPS + timestamp) | High (cryptographic sensor) | Fails to account for variable human judgment | High (for rigid tasks) |
5. Proposed Framework: Graduated Verification
Because no single mechanism can comprehensively solve the Last-Mile Oracle problem, relying on a binary trust model (where a transaction is either entirely trusted or entirely trustless) is a structural flaw. Thus why we propose a Graduated Verification Framework.
In physical engineering, a failure’s severity dictates its safety boundary. A bad API call is reversible; a severed water main is not. So an agent’s verification reqs need to scale dynamically with the financial stakes and/or the physical reversibility of the authorized action itself.
Low-Stakes / Reversible Transactions: For tasks with minimal financial value or those that can be easily undone a single automated signal is sufficient. Examples of this could be a digital micro task or a low value courier dropoff. Basic media attestation or a standard API webhook can clear this. As we need to optimize for speed and low friction.
Medium-Stakes Transactions: As financial value increases, the agent must require multi-factor physical verification to cross the cyber-physical boundary. This involves combining independent vectors: a cryptographically signed photograph, verified GPS telemetry from the worker’s device, and a time-bound cryptographic signature. This will work when spoofing these costs more then the value of the item/service itself.
High-Stakes / Irreversible Transactions: For high-value physical interventions (expensive home repairs, safety-critical mechanical work, physical asset movement, etc.), automated software attestation is mathematically insufficient. In a high stakes environment final settlement cannot and should not be released by the software agent alone. And it shall require human-in-the-loop confirmation to close the transaction.
To illustrate this interlock, consider a $2,000 main-line plumbing repair orchestrated by a home agent. The agent successfully executes the discovery and authorization phases, utilizing a standard like AP2 to lock the agreed-upon stablecoins into an escrow smart contract. The technician arrives, performs the physical labor, and submits a cryptographic proof-of-work via their device. However, rather than the agent automatically triggering the x402 settlement based on this potentially spoofable digital input, the contract enters a “Verification Halted” state. The homeowner is prompted to physically inspect the replaced pipe. Only after confirming the real-world outcome does the human user cryptographically sign the release of funds. In this scenario the agent handles sourcing, negotiating, and structuring the transaction, while the final authorization collapses to a single human attestation. This does not make the Last-Mile Oracle problem disappear; it relocates it from the paid technician, whose incentives are misaligned, to the homeowner, whose incentives are aligned with correct verification. We accept that tradeoff deliberately: the party best positioned to judge the physical outcome, and least incentivized to lie about it, becomes the root of trust for high-risk physical events.
6. Limitations and Methodology Note
This paper relies on a comparative architecture and threat-modeling methodology. Because the underlying protocols enabling agentic commerce (such as AP2, TAP, and x402) are nascent infrastructure introduced in 2025 and 2026, empirical, at-scale deployment data regarding physical failure modes does not yet exist. The scenarios presented are illustrative and designed to stress-test theoretical vulnerabilities.
The Deterministic Safety Boundary
In the pursuit of fully autonomous commerce, the industry often views human intervention as a failure of the system. However, when transitioning from digital to physical AI this becomes a mandatory architectural requirement.
An AI model may initiate a decision, but the system must determine whether that decision can safely become an action. When software attestation reaches its limit, trust will be rooted in hardware. In high-stakes agentic commerce, this takes the form of a physical interlock. Just as a hardware wallet requires a human to physically press a button to authorize a high-value on-chain transaction, the Last-Mile Oracle requires a physical hardware boundary to release high-stakes escrow.
To execute this we need to have the hardware device evolve from a static cryptographic key vault to an active element in the agentic workflow. We can do this by when an agent stages a high-stakes settlement the resulting smart contract interaction will require a hardware-backed signature to execute the final state change. The user provides an un-spoofable attestation by physically pressing a button on a trusted device.
Like mentioned before the agent orchestrates the logistics, negotiates the pricing, and stages the settlement. The final crossing into the physical boundary relies on a trusted execution environment and a human-in-the-loop. This paradigm shifts the role of the hardware wallet. It is no longer just protecting assets from digital theft; it is protecting the human user from their own autonomous agents. This ensures that while the economy becomes agentic, physical sovereignty remains deterministic. It leaves the final authority over physical reality strictly in human hands.
7. Conclusion
The agentic economy is rapidly standardizing how AI systems identify themselves, authorize spend, and settle payments. Without a verifiable mechanism to confirm physical-world outcomes these protocols will remain confined to digital-only transactions. The Last-Mile Oracle problem cannot be solved by better payment rails; it requires understanding the friction of the cyber-physical boundary. By implementing a Graduated Verification Framework, the industry can scale autonomous commerce safely. Matching the burden of proof to the physical stakes of the action and relying on hardware-anchored confirmation with human in the loop as the ultimate deterministic fallback. This will ensure that agents can operate at the speed of software without sacrificing the safety of the physical world.
- Eskandari, S., Salehi, M., Gu, W. C., & Clark, J. (2021). “SoK: Oracles from the Ground Truth to Market Manipulation.” 3rd ACM Conference on Advances in Financial Technologies (AFT ’21), September 26–28, 2021, Arlington, VA, USA. https://doi.org/10.1145/3479722.3480994
- Ethereum Foundation. (n.d.). ERC-8004: Trustless Agents Identity and Reputation Registries [Draft specification].
- Goenka, M. (2026). “TessPay: Verify-then-Pay Infrastructure for Trusted Agentic Commerce.” arXiv preprint. arXiv:2602.00213. https://arxiv.org/abs/2602.00213
- Hitachi Ventures. (2026). When Agents Enter the Physical World. Hitachi Ventures.
- Lin, Z., Wang, T., Shi, L., Zhang, S., & Cao, B. (2024). “Decentralized Physical Infrastructure Network (DePIN): Challenges and Opportunities.” arXiv preprint. https://doi.org/10.48550/arxiv.2406.02239
- Sarkar, D. (2023). “Generalised DePIN Protocol: A Framework for Decentralized Physical Infrastructure Networks.” arXiv preprint. https://doi.org/10.48550/arxiv.2311.00551
- Wang, Q. (2026). “When HTTP 402 Meets the Blockchain: Risks on Emerging x402 Payments.” arXiv preprint. arXiv:2607.19545. https://arxiv.org/abs/2607.19545
- Worldline & Google. (n.d.). Agent Payments Protocol (AP2) and the Mandate Model [Official documentation].
I hereby declare that this submission is my own original work. Where I have relied on the work, ideas, or data of others, I have properly cited and attributed those sources in accordance with academic standards. Artificial intelligence tools were utilized during the drafting process to assist with structural formatting, proofreading, and expanding upon the provided arguments; however, the core thesis, framework design, and conclusions remain solely my own.
Joey Kokinda · August 2026