测试大模型在多轮协作中处理私有信息的能力,发现其表现远低于潜力。
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
- 设计多轮协作游戏,评估模型在私有信息传递中的交互能力。
- 固定预算下增加对话轮次,模型仍无法超越非交互基线。
- 适合关注对话规划与信息管理的AI研究者参考。
我们提出一种可扩展且可验证的方法,通过一系列需要有效沟通私有信息的协作游戏,评估语言模型在多轮交互中的表现。该方法支持交互式规模分析:在固定令牌预算下,分配给不同轮次数量。结果发现,尽管存在巨大提升空间,语言模型在多轮协作中往往无法超越非交互基线(即一方先总结信息,另一方立即行动)的表现。这表明当前顶尖模型在规划和执行多轮协作对话方面仍存在显著缺陷。我们分析了对话中的语言特征,包括奉承性、信息密度和语篇连贯性。虽然没有单一语言因素能解释模型的协作弱点,但人类在完成类似任务时能以更优的令牌效率达成相似成功率,因其对话更具连贯性。主动管理私有信息是现实交流的核心特征,本工作旨在推动该能力的进一步发展。
原文摘要 · Abstract (English)
We present a scalable and verifiable methodology for evaluating language models in multi-turn interactions, using a suite of collaborative games that require effective communication about private information. This enables an interactive scaling analysis, in which a fixed token budget is divided over a variable number of turns. We find that language models often fail to use interactive collaboration to improve over the non-interactive baseline in which one agent summarizes its information and the other agent immediately acts, despite substantial headroom. This suggests that state-of-the-art models still suffer from significant weaknesses in planning and executing multi-turn collaborative conversations. We analyze the linguistic features of these dialogues, assessing the roles of sycophancy, information density, and discourse coherence. While there is no single linguistic explanation for the collaborative weaknesses of contemporary language models, we note that humans achieve comparable task success at superior token efficiency by producing more coherent dialogues. The proactive management of private information is a defining feature of real-world communication, and this work is designed to drive further progress on this capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。