让大模型组队答题,准确率最高提升20个百分点。
Can LLM Teams Play What? Where? When?

- 三类团队策略:投票、静默队长、对话队长
- 最佳团队准确率达44.23%,接近人类水平
- 模型分歧大时易出错,但交流能显著降低失误
大型语言模型在需要间接推理、文化知识和协作假设测试的任务上仍显不足。我们研究团队协作是否能提升模型在What? Where? When?(ChGK)这一强调集体推理的问答游戏中表现。引入三种团队策略:投票、静默团队(队长仅看最终答案)和话多团队(队长可查看答案与推理过程)。为避免数据泄露,使用2025年发布的572道ChGK题集进行评估。基于六种近期开源大模型,结果显示团队策略优于单模型基线,准确率最高提升20个百分点。最佳团队达44.23%准确率,接近已有数据中的人类团队表现。分析表明,模型间分歧强烈预示低准确率,但解释性沟通可显著缓解性能下降。进一步考察队长行为发现无自我偏好偏差;接触同伴推理过程能提升队长判断质量。整体来看,LLM团队主要起答案筛选与错误过滤作用,而非生成新解。研究强调交互重要性,提出自适应策略是多智能体系统的重要方向。
原文摘要 · Abstract (English)
Large language models (LLMs) remain limited on tasks requiring indirect reasoning, cultural knowledge, and coordinated hypothesis testing. We investigate whether team-based interaction improves LLM performance in What? Where? When? (ChGK), a quiz game designed to reward collective reasoning. We introduce three team strategies: Voting, Silent Team (the captain observes final answers), and Talkative Team (the captain observes both answers and rationales). To minimize data leakage, we evaluate these strategies on a dataset consisting of 572 ChGK questions released in 2025. Using six recent large-scale open models, we show that team-based strategies outperform single-model baselines, yielding gains of up to 20 percentage points in accuracy. The best team achieves 44.23% accuracy, and approaches human team performance on questions with available human statistics. Analysis of inter-model diversity reveals that disagreement strongly predicts lower accuracy, but explanatory communication substantially mitigates performance drops. We further examine captain behavior and find no evidence of self-preference bias; access to peer rationales improves captain judgments. Overall, LLM teams function primarily as answer selection and error-filtering mechanisms rather than generators of novel solutions. Our findings highlight the importance of interaction and suggest adaptive strategies as a promising direction for multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。