arXiv:2608.15530cs.CL2026-08

RLHF让摘要变中性,新方法揭示原因并缓解情感流失

Why Summaries Turn Neutral: Policy Attribution for Sentiment Drift in Reinforcement Learning from Human Feedback

论文配图:Why Summaries Turn Neutral: Policy Attribution for Sentiment Drift in Reinforcement Learning from Human Feedback
图 1 · 摘自论文原文
  • 通过梯度与逻辑分解追踪情感漂移源头
  • RLHF使摘要情感方差降低30-40%,跨语言普遍存在
  • 提出情感感知正则化,降低漂移18-22%且不损质量

基于人类反馈的强化学习(RLHF)虽提升大模型摘要的流畅性与安全性,却引发情感漂移:摘要过度中性化,失去情感特征。我们诊断出RL作为情感中和器的原因,提出策略归因(Policy Attribution)框架,利用梯度与对数几率分解,将漂移归因于奖励模型(RM)信号和KL惩罚。情感漂移源于在偏好不确定性下对“低风险”词汇的策略性偏好(Stiennon et al., 2020;Gao, Schulman, and Hilton, 2023)。在Reddit TL;DR与CNN/DailyMail数据集上,RLHF摘要得分更高,但情感方差下降30-40%。跨八种语言分析显示漂移具有语言无关性,形态更丰富的语言受抑制更明显(Krasitskii et al., 2026)。我们提出并验证了一种情感感知正则化技术,使漂移减少18-22%,同时保持摘要质量。代码与工具包将公开。

原文摘要 · Abstract (English)

Reinforcement learning with human feedback (RLHF) aligns LLMs with human preferences, improving summarization fluency and safety, but causes sentiment drift: overly neutral summaries stripped of emotional nuance. We diagnose why RL acts as a sentiment neutralizer and present Policy Attribution, a framework using gradient and logit decomposition to trace drift to reward model (RM) signals and KL (Kullback-Leibler) penalty. Sentiment drift reflects a strategic bias toward "low-risk" tokens maximizing expected rewards under preference uncertainty (Stiennon et al., 2020; Gao, Schulman, and Hilton, 2023). On Reddit TL;DR and CNN/DailyMail, RLHF summaries get higher rewards but show 30-40% lower sentiment variance. Cross-lingual analysis across eight languages shows language-independent drift, with morphologically richer languages more suppressed (Krasitskii et al., 2026). We propose and validate a sentiment-aware regularization technique reducing drift by 18-22% without harming summary quality. The code and toolkit will be public.

RLHF情感漂移大模型对齐策略归因

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。