arXiv:2606.08940cs.CL2026-06

发现强化学习生成摘要会削弱情感表达,提出改进方法保留情绪特征。

Multilingual Sentiment Aware Text Summarization A Reinforcement Learning Approach for Consistency Maintenance

论文配图:Multilingual Sentiment Aware Text Summarization A Reinforcement Learning Approach for Consistency Maintenance
图 1 · 摘自论文原文
  • 用强化学习优化摘要时,情感倾向易被弱化
  • 增加正则强度会加剧情感中性化,最高达37%的极性下降
  • 新方法针对情感词降低约束,适合多语言情感任务

基于人类反馈的强化学习(RLHF)显著提升了大模型在文本摘要中的质量与流畅性,但其对情感属性的影响尚不明确。本文研究了情感漂移现象:相比原文,RLHF生成的摘要系统性地趋向中性情感。我们在多个数据集、模型架构和八种语言上进行广泛实验,分析对齐目标如何影响情感保留。结果表明,情感漂移是普遍现象,且随KL正则化强度增加而加剧,说明对齐稳定性与情感真实性存在权衡。为此,我们提出策略归因框架,分解并量化RLHF目标各成分的贡献,发现KL正则化是所有设置下情感抑制的主要原因。基于此,我们设计了一种情感感知的KL正则化修改方案,有选择地放松对情感承载词的约束。实证结果表明,该方法有效缓解情感漂移,同时保持摘要质量。总体而言,当前对齐方法虽提升事实一致性和安全性,却可能无意中抑制情感表达,这促使开发显式考虑情感保留的对齐策略。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) has significantly improved the quality and fluency of large language models in text summarization. However, its impact on affective properties remains insufficiently understood. In this work, we study sentiment drift, a systematic shift toward neutral sentiment in RLHF-based summarization outputs compared to source texts. We conduct extensive experiments across multiple datasets, model architectures, and eight languages to analyze how alignment objectives influence sentiment preservation. Our results show that sentiment drift is a consistent phenomenon that becomes stronger with increased KL regularization strength, indicating a trade-off between alignment stability and affective fidelity. To explain this behavior, we introduce a Policy Attribution framework that decomposes the RLHF objective and quantifies the contribution of its components. Our analysis reveals that KL regularization is the primary driver of sentiment suppression across all settings. Based on these findings, we propose a sentiment-aware modification of the KL regularization term, which selectively reduces constraints on sentiment-bearing tokens. Empirical results demonstrate that this approach mitigates sentiment drift while maintaining summarization quality. Overall, our findings highlight a fundamental limitation of current alignment methods: while they improve factual consistency and safety, they may unintentionally suppress emotional expressiveness. This motivates the development of alignment strategies that explicitly account for affective preservation.

情感分析强化学习摘要生成多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。