arXiv:2604.20658cs.CLcs.CY2026-04

用行为经济学游戏预测大模型团队协作能力,提升科学任务表现

Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows

  • 通过六种经济博弈测试模型合作倾向,生成可量化合作画像
  • 能有效协调的模型在科学报告准确率、质量、完成度上均更优
  • 低成本诊断工具,适合需团队协作的大模型部署前筛选

由大型语言模型(LLMs)组成的多智能体系统正越来越多地用于协同科学推理与问题解决。这些系统需在共享资源(如GPU或信用额度)约束下协作,合作行为至关重要。行为经济学提供了多种隔离不同合作机制的博弈工具,但尚不清楚模型在这些简化场景中的行为是否能预测其在真实协作任务中的表现。本文对35个开源权重的LLM在六种行为经济学博弈中进行评测,结果表明,基于博弈生成的合作画像能稳健预测其在面向科学任务中的下游表现——在共享预算约束下,团队协作分析数据、构建模型并撰写科学报告。能够有效协调博弈、投入乘法式团队产出(而非采取利己策略)的模型,在准确性、质量和完成度三个指标上表现更优。该关联在控制多种因素后依然成立,表明合作倾向是独立于通用能力的可测量的LLM属性。因此,本研究提出的博弈框架可作为快速低成本的诊断工具,在昂贵的多智能体部署前筛选合作适配性。

原文摘要 · Abstract (English)

Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problem-solving. These systems require agents to coordinate under shared constraints, such as GPUs or credit balances, where cooperative behavior matters. Behavioral economics provides a rich toolkit of games that isolate distinct cooperation mechanisms, yet it remains unknown whether a model's behavior in these stylized settings predicts its performance in realistic collaborative tasks. Here, we benchmark 35 open-weight LLMs across six behavioral economics games and show that game-derived cooperative profiles robustly predict downstream performance in AI-for-Science tasks, where teams of LLM agents collaboratively analyze data, build models, and produce scientific reports under shared budget constraints. Models that effectively coordinate games and invest in multiplicative team production (rather than greedy strategies) produce better scientific reports across three outcomes, accuracy, quality, and completion. These associations hold after controlling for multiple factors, indicating that cooperative disposition is a distinct, measurable property of LLMs not reducible to general ability. Our behavioral games framework thus offers a fast and inexpensive diagnostic for screening cooperative fitness before costly multi-agent deployment.

多智能体合作评估科学计算行为博弈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。