arXiv:2607.23982cs.MAcs.AI2026-07

研究多智能体语言模型中的道德风险,发现信息共享依赖机制而非直接透露。

Moral Hazard in Multi-Agent Language Models

论文配图:Moral Hazard in Multi-Agent Language Models
图 1 · 摘自论文原文
  • 设计文本环境模拟团队协作中的隐藏行为问题,测试模型信息获取与沟通策略。
  • GEPA方法使团队成功率从22.2%提升至100%,查询使用率从51.1%降至0.3%。
  • 提出CREDIT算法,通过反事实重放优化机制,识别模型真实贡献瓶颈。

当社会有益的努力成本高、难以观察且主要惠及他人时,合作可能失败。基于霍姆斯特罗姆的团队道德风险模型,本文构建了对话道德风险游戏,将隐藏行动结构嵌入文本环境,让智能体在保留即时奖励与支付查询成本揭示隐藏安全事实之间抉择。评估了14个开源和4个前沿模型,涵盖信息获取、通信、下游使用与团队成功率。在每模型3,015次决策的匹配实验中,GPT-5.6 Sol、Claude Opus 4.8和Nemotron-3 Ultra在九种查询成本下接近理论私有份额边界,平均绝对误差分别为0.013、0.030、0.024;Muse Spark 1.1仅方向性响应,Fable 5持续高查询。不同训练方式(SFT、RLOO、SFT+RLOO、GEPA)引发异质性机制变化。其中,GEPA将团队成功率从22.2%提升至100.0%,同时将查询使用率从51.1%降至0.3%。冻结提示干预表明,成功依赖于学习到的排名标签映射,而非直接披露:改变该映射使成功率降至12.5%再至0.0%。本文提出CREDIT(反事实重放驱动的信息传递),一种对齐机制的多智能体提示优化算法,利用匹配的隐状态孪生体与全动作重放,奖励稳健因果贡献而非查询频率。在五个模型及多个随机种子下,CREDIT保持查询行为的同时揭示模型特定的信息获取与下游使用瓶颈。优化可通过对直接披露或学习有效信息结构达成相同总体结果,推动机制级评估与优化,而非仅关注团队成功。

原文摘要 · Abstract (English)

Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstrom's model of moral hazard in teams, the Dialogue Moral Hazard Game instantiates this hidden-action structure as a textual environment for language agents. An agent chooses between keeping an immediate local reward and paying a query cost to reveal a hidden safety fact that helps another agent's downstream decision. We evaluate fourteen open-weight and four frontier models using measures of information acquisition, communication, downstream use, and team success. In matched 3,015-decision-per-model experiments, GPT-5.6 Sol, Claude Opus 4.8, and Nemotron-3 Ultra track the derived private-share boundary across nine query costs, with mean absolute errors of 0.013, 0.030, and 0.024. Muse Spark 1.1 responds directionally, whereas Fable 5 remains query-saturated. SFT, RLOO, SFT+RLOO, and GEPA produce heterogeneous mechanism changes. GEPA raises Muse team success from 22.2% to 100.0% while reducing query use from 51.1% to 0.3%. Frozen-prompt interventions show that this success depends on a learned rank-label mapping rather than direct revelation: changing the mapping reduces team success from 100.0% to 12.5% and then 0.0%. We introduce CREDIT (Counterfactual Replay for Evidence-Driven Information Transfer), a mechanism-aligned multi-agent prompt-optimization algorithm that uses matched hidden-state twins and total-action replay to reward robust causal contribution rather than query frequency. Across five models and multiple seeds, CREDIT preserves query-mediated behavior while revealing model-specific acquisition and downstream-use bottlenecks. Optimization can reach the same aggregate outcome through direct revelation or a learned effective information structure, motivating mechanism-level evaluation and optimization rather than team success alone.

多智能体道德风险机制优化信息传递

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。