arXiv:2509.09071cs.AIcs.GT2025-09被引 12

比较人类与大模型在多方讨价还价中的表现差异。

Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining

  • 用相同条件测试人类、大模型和贝叶斯代理的博弈策略。
  • 大模型倾向保守让步,成功率高;人类更讲公平,但易被拒。
  • 性能相近却策略迥异,揭示评估需关注行为机制差异。

市场日益接纳大语言模型(LLMs)作为自主决策代理。随着这一转变,评估这些代理相对于其人类及任务特定统计前辈的行为变得至关重要。本文通过一项实证研究,比较了216名人类、多个前沿大模型和定制贝叶斯代理在相同条件下参与动态多玩家讨价还价游戏的表现。贝叶斯代理通过激进的交易提案获得最高总收益,尽管常被拒绝。人类与大模型在其群体内达到相当的总收益,但交易策略不同:大模型偏好保守、让步型提案,通常被其他大模型接受;而人类提出的方案符合公平规范,但更易被拒绝。结果表明,性能一致——代理评估中的常见基准——可能掩盖了大模型在复杂多代理交互中实质性程序差异。

原文摘要 · Abstract (English)

Markets increasingly accommodate large language models (LLMs) as autonomous decision-making agents. As this transition occurs, it becomes critical to evaluate how these agents behave relative to their human and task-specific statistical predecessors. In this work, we present results from an empirical study comparing humans (N=216), multiple frontier LLMs, and customized Bayesian agents in dynamic multi-player bargaining games under identical conditions. Bayesian agents extract the highest surplus with aggressive trade proposals that are frequently rejected. Humans and LLMs achieve comparable aggregate surplus within their groups, but exhibit different trading strategies. LLMs favor conservative, concessionary proposals that are usually accepted by other LLMs, while humans propose trades that are consistent with fairness norms but are more likely to be rejected. These findings highlight that performance parity -- a common benchmark in agent evaluation -- can mask substantive procedural differences in how LLMs behave in complex multi-agent interactions.

多智能体博弈大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。