推理未必提升角色扮演能力,反而可能降低表现。
Reasoning Does Not Necessarily Improve Role-Playing Ability
- 对比直接扮演、思维链和优化推理模型三种策略
- 思维链导致角色扮演性能下降,大模型仍不擅长高级扮演
- 适合研究角色扮演优化与强化学习的学者
角色扮演类大语言模型在学术与商业领域应用日益广泛,对高精度角色扮演模型的需求不断增长。与此同时,推理技术的快速发展持续推动大模型性能边界。这一实践需求与技术演进的交汇提出关键问题:推理技术能否提升大模型的角色扮演能力?为此,我们基于6个角色扮演基准、24个大模型及3种不同扮演策略,系统比较了直接零样本扮演、思维链(CoT)扮演以及推理优化模型扮演的效果。结果表明:思维链可能降低角色扮演表现,推理优化模型不适用于角色扮演,推理能力破坏角色扮演的规模效应,大模型在高级角色扮演上仍显不足,且中文角色扮演表现优于英文。基于实验结果,我们提出两个未来方向:面向角色感知的思维链与基于强化学习的角色扮演,以提升模型在真实场景中的适应性、一致性和有效性。
原文摘要 · Abstract (English)
The application of role-playing large language models (LLMs) is rapidly expanding in both academic and commercial domains, driving an increasing demand for high-precision role-playing models. Simultaneously, the rapid advancement of reasoning techniques has continuously pushed the performance boundaries of LLMs. This intersection of practical role-playing demands and evolving reasoning capabilities raises an important research question: "Can reasoning techniques enhance the role-playing capabilities of LLMs?" To address this, we conduct a comprehensive study using 6 role-playing benchmarks, 24 LLMs, and 3 distinct role-playing strategies, comparing the effectiveness of direct zero-shot role-playing, role-playing with Chain-of-Thought (CoT), and role-playing using reasoning-optimized LLMs. Our findings reveal that CoT may reduce role-playing performance, reasoning-optimized LLMs are unsuitable for role-playing, reasoning ability disrupts the role-playing scaling law, large models still lack proficiency in advanced role-playing, and Chinese role-playing performance surpasses English role-playing performance. Furthermore, based on extensive experimental results, we propose two promising future research directions: Role-aware CoT for improving role-playing LLMs and Reinforcement Learning for role-playing LLMs, aiming to enhance the adaptability, consistency, and effectiveness of role-playing LLMs for both research and real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。