用推理游戏训练大模型多视角协作,提升复杂场景下的决策能力。
Who is Undercover? Guiding LLMs to Explore Multi-Perspective Team Tactic in the Game
- 通过轮流发言与投票,让大模型模拟多角度思考与自我认知。
- 在游戏测试中,模型展现出策略性隐瞒与沟通的理性决策能力。
- 适合研究群体协作、公平决策的大模型应用者参考。
大型语言模型(LLMs)在复杂任务中扮演关键角色,但在开放决策问题上仍面临挑战。为此,我们以语言逻辑游戏《谁是卧底?》(WIU)为实验平台,提出多视角团队策略(MPTT)框架。该框架旨在培养大模型的人类式语言表达逻辑、多维思维及复杂情境中的自我认知。通过交替进行发言与投票环节,融合自视角、身份判定、自我反思、自我总结和多轮找队友等技术,大模型代理在策略性隐藏与交流中做出理性决策,促进类人信任关系的形成。初步结果显示,结合WIU的MPTT能够利用大模型的认知能力,构建可模拟真实社会的决策框架,有助于少数群体的沟通与表达,推动决策中的公平与多样性。此外,人机协同实验表明,大模型可通过交互学习并适应人类行为,具备在社会决策中主动参与的潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are pivotal AI agents in complex tasks but still face challenges in open decision-making problems within complex scenarios. To address this, we use the language logic game ``Who is Undercover?'' (WIU) as an experimental platform to propose the Multi-Perspective Team Tactic (MPTT) framework. MPTT aims to cultivate LLMs' human-like language expression logic, multi-dimensional thinking, and self-perception in complex scenarios. By alternating speaking and voting sessions, integrating techniques like self-perspective, identity-determination, self-reflection, self-summary and multi-round find-teammates, LLM agents make rational decisions through strategic concealment and communication, fostering human-like trust. Preliminary results show that MPTT, combined with WIU, leverages LLMs' cognitive capabilities to create a decision-making framework that can simulate real society. This framework aids minority groups in communication and expression, promoting fairness and diversity in decision-making. Additionally, our Human-in-the-loop experiments demonstrate that LLMs can learn and align with human behaviors through interactive, indicating their potential for active participation in societal decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。