arXiv:2608.01341cs.AI2026-08

为自动代理设计可自适应的支付决策层,实现有限钱包下的智能服务选购。

402Pilot: An x402 Decision Layer for Autonomous Agent Micropayments

论文配图:402Pilot: An x402 Decision Layer for Autonomous Agent Micropayments
图 1 · 摘自论文原文
  • 构建协议无关的购买决策层,根据钱包压力动态选服务商。
  • 在固定预算下仅用39%~43%资金,仍保持优质服务与市场变化响应能力。
  • 适合需要长期运行、资源受限的自动化系统开发者使用。

可编程支付协议如x402支持按请求微支付,但无法决定自主代理在有限钱包下应购买哪个可付费服务。我们提出将此问题建模为代理原生支付决策:在钱包压力、仅付款后反馈、市场条件变化的背景下进行上下文相关服务选择。为此,我们设计402Pilot,一个位于自主代理与支付执行之间的协议无关购买决策层,实现选择可付费服务的策略。我们以PA-DCT为例,这是一种考虑支付的折扣上下文汤普森采样策略,在钱包压力下持续学习并调整采购决策。为评估买家侧支付策略,我们引入402Pilot-Bench,一个包含823个任务、五个异构服务管道和三种市场模式的冻结回放基准,每种配置均在30组种子上评估。PA-DCT在非预言机策略中表现出最优的固定钱包自适应权衡:在仅花费39%至43%预算的同时保持竞争力的服务质量,并能随市场变化重新分配支出。其在价格冲击下取得最佳非预言机PA-gap/T表现,在质量、投资回报率和PA-gap/T的九种情景-指标组合中整体排名最优且最差情况也表现良好。与学习基线对比及组件消融实验进一步验证了该策略的有效性与设计合理性。结果表明,可编程支付需辅以能够学习服务价值并相应调整采购决策的买方侧决策能力。

原文摘要 · Abstract (English)

Programmable-payment protocols such as x402 enable per-request micropayments, but they do not determine which payable service an autonomous agent should buy under a finite wallet. We formulate this buyer-side problem as agent-native payment decision-making: contextual provider selection under wallet pressure, chosen-only paid feedback, and changing market conditions. We propose 402Pilot, a protocol-agnostic buyer-side decision layer between autonomous agents and payment execution that implements purchasing policies for selecting among payable providers. We instantiate it with PA-DCT, a payment-aware discounted contextual Thompson-sampling policy that adapts purchasing decisions under wallet pressure while learning from post-payment feedback. To evaluate buyer-side payment policies, we introduce 402Pilot-Bench, a frozen-replay benchmark spanning 823 tasks, five heterogeneous provider pipelines, and three market regimes, each evaluated over 30 paired seeds. PA-DCT achieves the strongest fixed-wallet adaptive trade-off among non-oracle policies: it maintains competitive service quality while spending only 39 to 43 percent of the wallet and reallocates spending as market conditions change. It attains the best non-oracle PA-gap/T under the price shock and the best mean and worst-case ranks across the nine scenario-metric combinations of quality, ROI, and PA-gap/T. Comparisons with learning baselines and component ablations further support the effectiveness and design of the proposed decision policy. These results suggest that programmable payment must be complemented by buyer-side decision-making capable of learning service value and adapting purchasing decisions accordingly.

支付协议自主代理决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。