arXiv:2501.01652cs.CL2025-01ACL被引 4

用推理游戏评测大模型在复杂社交环境中的表现。

MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments

  • 设计八套剧情剧本,模拟多样社交互动场景。
  • 提出四项评估指标,量化信任、线索调查等能力。
  • 发现GPT-4在复杂角色互动中仍存明显短板。

大型语言模型(LLMs)在环境感知、基于推理的决策和模拟复杂人类行为方面展现出显著能力,尤其在交互式角色扮演情境中。本文提出多宇宙互动角色扮演能力综合评估框架MIRAGE,用于评估LLMs在谋杀谜案游戏中的高级人类行为表现。MIRAGE包含八套精心设计的剧本,涵盖多样主题与风格,提供丰富仿真环境。为评估模型表现,MIRAGE采用四种方法:信任倾向指数(TII)衡量信任与怀疑动态变化,线索调查能力(CIC)评估信息搜集能力,互动能力指数(ICI)评估角色扮演水平,脚本合规指数(SCI)评估对指令的理解与遵循能力。实验表明,即使主流模型如GPT-4在面对MIRAGE的复杂性时也面临显著挑战。数据集与仿真代码已开源于github。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly in interactive role-playing contexts. This paper introduces the Multiverse Interactive Role-play Ability General Evaluation (MIRAGE), a comprehensive framework designed to assess LLMs' proficiency in portraying advanced human behaviors through murder mystery games. MIRAGE features eight intricately crafted scripts encompassing diverse themes and styles, providing a rich simulation. To evaluate LLMs' performance, MIRAGE employs four distinct methods: the Trust Inclination Index (TII) to measure dynamics of trust and suspicion, the Clue Investigation Capability (CIC) to measure LLMs' capability of conducting information, the Interactivity Capability Index (ICI) to assess role-playing capabilities and the Script Compliance Index (SCI) to assess LLMs' capability of understanding and following instructions. Our experiments indicate that even popular models like GPT-4 face significant challenges in navigating the complexities presented by the MIRAGE. The datasets and simulation codes are available in \href{https://github.com/lime728/MIRAGE}{github}.

角色扮演社交推理评估框架大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。