arXiv:2601.04716cs.CL2026-01中稿 · EMNLP被引 1

发现角色扮演中道德属性是性能瓶颈,提出无需训练的修复方法。

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

  • 拆解角色设定为熟悉度、结构、倾向三轴,系统测试影响
  • 道德倾向差导致性能显著下降,且越对齐越严重
  • 提出无训练的对比解码法,有效缓解道德偏差问题

尽管大语言模型角色扮演代理发展迅速,但尚不清楚哪些角色设定要素真正影响表现质量。为此,我们提出一种系统的诊断框架,将角色设定沿三个维度解耦:熟悉度(已知 vs. 未知)、结构(结构化 vs. 非结构化)和倾向(道德 vs. 非道德)。采用统一的层级架构(5个维度,28个字段),构建包含211个角色的受控数据集,并在单轮与多轮交互中评估五种LLM。结果揭示显著不对称性:熟悉度与结构影响微乎其微,而道德倾向对非道德角色产生大规模且一致的性能下降。进一步分析表明,道德-非道德差距在后SFT对齐阶段被放大,且不同属性间的退化程度差异显著。为缓解此瓶颈,我们提出场感知对比解码(FACD),一种无需训练的策略,能增强被抑制的倾向敏感信号,显著缩小性能差距,同时不损害道德角色的表现。

原文摘要 · Abstract (English)

While Large Language Model (LLM) role-playing agents have advanced rapidly, it remains unclear which profile elements genuinely drive role-playing quality. To bridge this gap, we introduce a systematic diagnostic framework that disentangles the impact of character profiles along three axes: Familiarity (Known vs. Unknown), Structure (Structured vs. Unstructured), and Disposition (Moral vs. Immoral). Utilizing a unified hierarchical schema (5 dimensions, 28 fields), we construct a controlled dataset of 211 personas and evaluate five LLMs on both single- and multi-turn interactions. Our results reveal a striking asymmetry: \textbf{Familiarity} and \textbf{Structure} show negligible impact, while \textbf{Disposition} produces large, consistent performance degradation for immoral characters across all conditions. Further analyses suggest that the Moral--Immoral gap is amplified by post-SFT alignment, and that this degradation varies substantially across profile attributes. To mitigate this bottleneck, we propose Field-Aware Contrastive Decoding (FACD), a training-free strategy that amplifies suppressed disposition-sensitive signals, significantly closing the performance gap without sacrificing moral-character performance.

角色扮演语言模型对齐偏差解耦分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。