让大模型像人一样思考角色,提升对话真实感
HER: Human-like Reasoning and Reinforcement Learning for LLM Role-playing
- 区分第一人称推理与第三人称分析,模拟角色内心活动
- 在CoSER上提升30.26分,在最小最大对弈任务上提升14.97%
- 适合需要深度角色扮演的应用,如虚拟陪伴和游戏
大语言模型角色扮演已成为陪伴、内容创作和数字游戏等应用的关键能力。尽管现有模型能较好捕捉人物语气与知识,但对其行为背后内在思维的模拟仍存在挑战。现有方法主要存在两大缺陷:缺乏高质量的推理轨迹数据,以及与人类偏好不一致的奖励信号。本文提出HER框架,实现认知级角色模拟。HER引入双层思维机制,区分角色的第一人称思考与模型的第三人称分析。通过逆向工程构建带推理标注的角色扮演数据,并建立符合人类偏好的原则与奖励模型。基于Qwen3-32B,结合监督与强化学习训练HER模型。大量实验验证了该方法的有效性:模型显著优于Qwen3-32B基线,在CoSER基准上提升30.26分,在最小最大角色扮演测试中提升14.97%。相关数据集、原则与模型已公开,以推动后续研究。
原文摘要 · Abstract (English)
LLM role-playing, i.e., using LLMs to simulate specific personas, has emerged as a key capability in various applications, such as companionship, content creation and digital games. While current models effectively capture character tones and knowledge, simulating the inner thoughts behind their behaviors remains a challenge. Towards cognitive simulation in LLM role-play, previous efforts mainly suffer from two deficiencies: lacking data with high-quality reasoning traces, and lacking reliable reward signals aligned with human preferences. In this paper, we propose HER, a unified framework for cognitive-level persona simulation. HER introduces dual-layer thinking, which distinguishes characters' first-person thinking from LLMs' third-person thinking. To bridge these gaps, we curate reasoning-augmented role-playing data via reverse engineering, and construct human-aligned principles and reward models. Leveraging these resources, we train HER models based on Qwen3-32B via supervised and reinforcement learning. Extensive experiments validate the effectiveness of our approach. Notably, our models significantly outperform the Qwen3-32B baseline, achieving a 30.26 improvement on the CoSER benchmark and a 14.97% gain on the Minimax Role-Play Bench. Our datasets, principles, and models are released to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。