arXiv:2601.22184cs.GTcs.LG2026-01被引 2

大模型无需沟通也能协作,但依赖特定焦点,且缺乏文化常识。

Tacit Coordination of Large Language Models

  • 通过焦点策略让大模型在无通信下实现高效协作
  • 20多个模型中,多数表现媲美甚至超过人类
  • 适合人机协同场景,但需警惕文化认知差异

大型语言模型(LLMs)越来越多地应用于无需通信的多智能体协作场景,从人机交互到安全关键任务。人类常借助焦点(focal points)——对所有参与者都显眼的解决方案——克服沟通缺失。我们首次大规模评估了大模型在合作与竞争游戏中焦点出现的时机与原因,涵盖真实搜救场景,揭示焦点如何促成有效协作。在超过20个开源与闭源模型中,大模型展现出惊人的无沟通协调能力,通常可媲美或超越人类。然而,这些模型在需要数值常识或文化敏感性判断的任务中持续失败。我们还测试了无需学习的简单策略,显著提升了大模型之间及人机间的协作效果。结果表明,现代大模型具有强大的协调潜力,但也存在社会认知局限,其内在的显眼性(salience)表征与人类不同。研究警示:部署大模型进行协作时,不应默认其具备人类的文化与感知基础。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in multi-agent settings that require coordination without communication, from human-AI interaction to safety-critical scenarios. Humans often overcome the absence of communication through focal points: salient solutions that naturally stand out to all participants. We present the first large-scale evaluation of how, when, and why focal points emerge in LLMs, comparing their behaviour with humans across cooperative and competitive games, including realistic search and rescue scenarios, demonstrating when focal points enable effective coordination. Across more than 20 open- and closed-source models, we find that LLMs exhibit a remarkable ability to coordinate without communication, often matching or outperforming humans. However, the same models consistently fail in tasks requiring numerical common sense or culturally nuanced notions of salience. We additionally evaluate simple learning-free strategies that substantially improve coordination both among LLMs and between humans and LLMs. Our results reveal striking coordination capabilities, as well as social limitations in modern LLMs, and offer new insight into the latent notions of salience encoded within them. Our findings caution against assuming that LLMs share humans' cultural and perceptual substrate when deployed in coordination settings.

大模型协作无通信焦点策略人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。