arXiv:2602.06008cs.AIcs.LG2026-02被引 13

用语言谈判的多智能体系统,模拟真实买卖交易场景。

AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions

  • 构建语言驱动的多智能体谈判框架,支持多方交互
  • 110+任务覆盖双边到多对多市场,评估达成率与效率
  • 揭示大模型长期策略推理短板,适合研究智能交易

大型语言模型(LLM)代理正被期待自主进行谈判、协调与交易,但现有基准缺乏对多智能体间语言媒介经济互动的系统性评估。我们提出AgenticPay,一个基于自然语言的多智能体买家-卖家谈判基准与仿真框架。该框架模拟买家与卖家拥有私有约束和依赖产品的估值,在多轮语言协商中达成协议,而非仅依赖数值竞价。系统支持超过110项任务,涵盖双边谈判到多对多市场,具备结构化动作提取与可行性、效率、福利等评估指标。对主流专有及开源大模型的基准测试显示其在谈判表现上存在显著差距,暴露出长周期战略推理能力不足的问题,确立了AgenticPay作为研究智能商业与语言化市场交互的基础。代码与数据集见:https://github.com/SafeRL-Lab/AgenticPay。

原文摘要 · Abstract (English)

Large language model (LLM)-based agents are increasingly expected to negotiate, coordinate, and transact autonomously, yet existing benchmarks lack principled settings for evaluating language-mediated economic interaction among multiple agents. We introduce AgenticPay, a benchmark and simulation framework for multi-agent buyer-seller negotiation driven by natural language. AgenticPay models markets in which buyers and sellers possess private constraints and product-dependent valuations, and must reach agreements through multi-round linguistic negotiation rather than numeric bidding alone. The framework supports a diverse suite of over 110 tasks ranging from bilateral bargaining to many-to-many markets, with structured action extraction and metrics for feasibility, efficiency, and welfare. Benchmarking state-of-the-art proprietary and open-weight LLMs reveals substantial gaps in negotiation performance and highlights challenges in long-horizon strategic reasoning, establishing AgenticPay as a foundation for studying agentic commerce and language-based market interaction. Code and dataset are available at the link: https://github.com/SafeRL-Lab/AgenticPay.

多智能体语言谈判经济仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。