研究大模型代理在信息不对称下的协作沟通与验证机制。
Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry
- 设计桌游场景模拟信息不对称的协作任务。
- 有验证器的代理任务完成率提升,理解更深入。
- 适合关注AI协作可信性与可解释性的研究者。
尽管大型语言模型(LLM)代理常被用于基于语言描述的目标规划与执行,但其协同完成联合目标的能力尚未得到充分探索。本文聚焦于信息不对称条件下的任务协作,即代理间知识与技能存在差异,需合作完成共同任务。为此,将经典的爱因斯坦谜题扩展为桌面游戏,要求两个LLM代理通过推理、通信与行动满足空间和关系约束以解谜。采用微调+验证器框架,使代理具备多种沟通策略及来自环境的验证信号。实证结果表明,对齐的沟通至关重要,尤其当代理兼具信息获取与提供能力时。有趣的是,无通信的代理仍能实现高任务性能,但分析显示其缺乏真正规则理解,且人类评估者信任度较低。通过引入基于环境的验证器,显著提升了代理对任务规则的理解与任务完成能力,推动了更安全、可解释的AI协作。
原文摘要 · Abstract (English)
While Large Language Model (LLM) agents are often approached from the angle of action planning/generation to accomplish a goal (e.g., given by language descriptions), their abilities to collaborate with each other to achieve a joint goal are not well explored. To address this limitation, this paper studies LLM agents in task collaboration, particularly under the condition of information asymmetry, where agents have disparities in their knowledge and skills and need to work together to complete a shared task. We extend Einstein Puzzles, a classical symbolic puzzle, to a table-top game. In this game, two LLM agents must reason, communicate, and act to satisfy spatial and relational constraints required to solve the puzzle. We apply a fine-tuning-plus-verifier framework in which LLM agents are equipped with various communication strategies and verification signals from the environment. Empirical results highlight the critical importance of aligned communication, especially when agents possess both information-seeking and -providing capabilities. Interestingly, agents without communication can still achieve high task performance; however, further analysis reveals a lack of true rule understanding and lower trust from human evaluators. Instead, by integrating an environment-based verifier, we enhance agents' ability to comprehend task rules and complete tasks, promoting both safer and more interpretable collaboration in AI systems. https://github.com/Roihn/EinsteinPuzzles
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。