arXiv:2503.08193cs.AI2025-03EMNLP被引 10

构建角色内省推理基准,提升语言代理的内心逻辑生成能力

Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents

  • 提出通过记忆检索、反应预测与动机合成生成角色内心思考
  • 在黄金集和白银集上均超越现有方法,验证内省推理有效性
  • 适合研究角色扮演智能体、对话系统与认知建模的学者使用

基于大语言模型的角色扮演语言代理(RPLAs)在各类应用中备受关注。尽管链式思维在多项任务中展现价值,但RPLAs的内部思考过程仍缺乏探索。理解角色的内心想法对发展高级角色扮演代理至关重要。本文提出ROLETHINK,一个基于文学作品构建的新型基准,用于评估角色思想生成能力。我们定义了内省推理任务,包含两部分:黄金集(将生成的思想与原角色独白对比)和银色集(以专家合成的角色分析为参考)。为应对挑战,我们提出MIRROR方法,通过检索记忆、预测角色反应并合成动机来生成角色思想。大量实验表明,内省推理对RPLAs极为重要,且MIRROR始终优于现有方法。资源已开源:https://github.com/airaer1998/RPA_Thought。

原文摘要 · Abstract (English)

Recent advances in LLM-based role-playing language agents (RPLAs) have attracted broad attention in various applications. While chain-of-thought reasoning has shown importance in many tasks for LLMs, the internal thinking processes of RPLAs remain unexplored. Understanding characters' inner thoughts is crucial for developing advanced RPLAs. In this paper, we introduce ROLETHINK, a novel benchmark constructed from literature for evaluating character thought generation. We propose the task of inner thought reasoning, which includes two sets: the gold set that compares generated thoughts with original character monologues, and the silver set that uses expert synthesized character analyses as references. To address this challenge, we propose MIRROR, a chain-of-thought approach that generates character thoughts by retrieving memories, predicting character reactions, and synthesizing motivations. Through extensive experiments, we demonstrate the importance of inner thought reasoning for RPLAs, and MIRROR consistently outperforms existing methods. Resources are available at https://github.com/airaer1998/RPA_Thought.

角色扮演内省推理链式思维语言代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。