测试大模型在博弈中的策略协调能力,模拟真实贸易场景。
Cattle Trade: A Multi-Agent Benchmark for LLM Bluffing, Bidding, and Bargaining

- 构建多智能体交易环境,融合拍卖、虚报、讨价还价等复杂行为。
- 7个低成本语言模型表现不及2个启发式代码代理,关键在于资源管理与策略适应。
- 揭示大模型常见失败模式:过度出价、自买自卖、破产发起挑战等。
我们提出Cattle Trade,一个用于评估大语言模型(LLMs)在不完全信息、对抗性交互和资源约束下的战略推理能力的多智能体基准。该基准将拍卖、隐藏报价交易挑战(TCs)、讨价还价、虚报、对手建模和资源分配整合在一个持续50至60回合的长周期游戏中。不同于以往仅孤立测试某项能力的基准,Cattle Trade 考察智能体是否能在具有冲突激励的竞争性经济游戏中协同运用多种能力。系统记录每次出价、TC报价、反报价及卡片选择,支持超越最终得分或胜率的行为分析。我们在242场游戏中评估了7个成本高效的语言模型和3个确定性代码代理。结果显示,战略连贯性——尤其是支出效率、资源自律和阶段适配出价——比总支出或单一子技能更显著影响排名。两个启发式代码代理优于大多数测试过的LLMs。行为轨迹揭示了大模型的重复性失败模式,包括过度出价、自我出价、破产发起交易挑战以及对对手状态适应能力弱。评估智能体能力需依赖能检验多种能力联合部署的基准,且须置于存在冲突激励、不确定性与经济动态的多智能体环境中。
原文摘要 · Abstract (English)
We introduce \textsc{Cattle Trade, a multi-agent benchmark for evaluating large language models (LLMs) as agents in strategic reasoning under imperfect information, adversarial interaction, and resource constraints. The benchmark combines auctions, hidden-offer trade challenges (TCs), bargaining, bluffing, opponent modeling, and resource allocation within a single long-horizon game lasting 50--60 turns. Unlike prior agent benchmarks that test these abilities in isolation, \textsc{Cattle Trade} evaluates whether agents integrate them across a competitive, multi-agent economic game with conflicting incentives. The benchmark logs every bid, TC offer, counteroffer, and card selection, enabling behavioural analysis beyond final scores or win rates. We evaluate seven cost-efficient language models and three deterministic code agents across 242 games. Strategic coherence, in particular spending efficiency, resource discipline, and phase-adaptive bidding, is associated with rank more strongly than spending volume or any single subskill. Two heuristic code agents outperform most tested LLMs, and behavioural traces surface recurring LLM failure modes including overbidding, self-bidding, bankrupt TC initiation, and weak opponent-state adaptation. Evaluating agentic competence requires benchmarks that test the joint deployment of multiple capabilities in multi-agent environments with conflicting incentives, uncertainty, and economic dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。