构建开放正和谈判环境,评估大模型在复杂经济互动中的表现
SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game

- 设计三阶段动态博弈框架:语言协商、生产转换、密封投标
- 强模型获更高收益,但仍存在资源误判与长期规划不足
- 适合研究多智能体协作、经济推理与大模型决策能力的学者
评估大语言模型代理需要超越静态推理和零和博弈的动态环境。现实经济互动常为开放式且混合动机:智能体需协商、创造正和收益、争夺稀缺资产,并在延迟回报下规划。我们提出 SidConArena,一个用于评估大模型代理在开放式正和谈判中的新基准框架。该框架将多智能体经济建模为有限时域部分可观测随机博弈,包含三个耦合阶段:带绑定交易的语言协商、基于确定性转换器的生产,以及长期资产的密封投标拍卖。框架结合结构化观测、阶段感知的代理调度、神经符号动作接口与异步执行,实现自由交互的同时保证规则可验证评估。在同质与异质锦标赛中,前沿模型取得更高经济成果,但智能体仍存在资源误判、被动谈判及长周期投资规划能力受限等问题。
原文摘要 · Abstract (English)
Evaluating LLM agents requires dynamic environments that go beyond static reasoning and zero-sum games. Real-world economic interaction is often open-ended and mixed-motive: agents must negotiate, create positive-sum surplus, compete for scarce assets, and plan under delayed returns. We introduce SidConArena, a new benchmark framework for evaluating LLM agents in open-ended, positive-sum bargaining. SidConArena formalizes a multi-player economy as a finite-horizon partially observable stochastic game with three coupled phases: natural-language negotiation with binding trades, deterministic converter-based production, and sealed-bid auctions for long-term assets. The framework combines structured observations, phase-aware agent dispatching, a neural-symbolic action interface, and asynchronous execution, enabling free-form interaction while preserving rule-grounded evaluation. Across homogeneous and heterogeneous tournaments, stronger frontier models achieve higher economic outcomes, yet agents still misvalue resources, bargain passively, and remain limited in long-horizon investment planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。