用信息论分析角色扮演模型泛化失效原因,找到关键风险点。
Understanding Generalization in Role-Playing Models via Information Theory
- 提出信息论指标R-EMID,量化角色、用户、对话变化对性能影响。
- 发现用户分布偏移是导致性能下降的最大风险,占比最高。
- 结合强化学习优化对话生成概率估计,提升模型适应能力。
角色扮演模型(RPMs)在实际应用中广泛使用,但在真实场景部署时性能下降明显,这主要源于用户、角色及对话组合的分布偏移。现有方法如大模型作为裁判(LLM-as-a-judge)难以提供细粒度诊断,缺乏系统性框架来刻画RPM泛化行为。为此,本文提出一种基于推理的有效互信息差异(R-EMID)指标,以可解释方式衡量性能退化。同时推导出R-EMID的上界,用于预测最差情况下的泛化表现,并理论揭示各类分布偏移对性能的影响机制。此外,提出一种协同进化强化学习框架,自适应建模用户、角色与对话上下文间的关联,从而更准确估计对话响应生成概率,这对计算R-EMID至关重要。最终通过R-EMID评估多种RPM的泛化性能,结果表明:用户分布偏移带来的风险最高,而强化学习是提升泛化能力最有效的手段。
原文摘要 · Abstract (English)
Role-playing models (RPMs) are widely used in real-world applications but underperform when deployed in the wild. This degradation can be attributed to distribution shifts, including user, character, and dialogue compositional shifts. Existing methods like LLM-as-a-judge fall short in providing a fine-grained diagnosis of how these shifts affect RPM generalization, and thus there lack formal frameworks to characterize RPM generalization behaviors. To bridge these gaps, we introduce an information-theoretic metric, named reasoning-based effective mutual information difference (R-EMID), to measure RPM performance degradation in an interpretable way. We also derive an upper bound on R-EMID to predict the worst-case generalization performance of RPMs and theoretically reveal how various shifts contribute to the RPM performance degradation. Moreover, we propose a co-evolving reinforcement learning framework to adaptively model the connection among user, character, and dialogue context and thus enhance the estimation of dialogue response generation probability, which is critical for calculating R-EMID. Finally, we evaluate the generalization performance of various RPMs using R-EMID, finding that user shift poses the highest risk among all shifts and reinforcement learning is the most effective approach for enhancing RPM generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。