arXiv:2605.26840cs.CL2026-05EMNLP被引 2

用多个不完美指标的偏好学习,提升摘要事实一致性。

Optimising Factual Consistency in Summarisation via Preference Learning from Multiple Imperfect Metrics

论文配图:Optimising Factual Consistency in Summarisation via Preference Learning from Multiple Imperfect Metrics
图 1 · 摘自论文原文
  • 通过组合多个弱性指标得分生成偏好数据,避免复杂奖励设计
  • 在不同解码策略下生成相似摘要对,捕捉细微事实差异
  • 无需人工标注,小模型也能达到大模型的事实一致性水平

强化学习中常以评估指标作为奖励来提升语言模型特定能力。然而,对于事实一致性的摘要任务,现有指标仍不完善,难以有效引导模型行为。尽管单一事实性指标不可靠,但其组合能更全面捕捉多样化的事实错误。本文提出一种自动化训练流程,通过聚合多个弱指标得分来提升摘要的事实一致性。该方法将指标得分映射为偏好,并过滤掉指标间高分歧的样本。针对每个源文档,通过改变解码策略生成语义相近的摘要对,使模型能够从细微词汇差异引发的事实差异中学习。整个过程仅依赖源文档构建高质量偏好数据集。实验表明,该方法在多种模型上均实现持续的事实一致性提升,涵盖早期编码器-解码器架构到现代大语言模型,小模型的表现可媲美大模型。

原文摘要 · Abstract (English)

Reinforcement learning with evaluation metrics as rewards is widely used to enhance specific capabilities of language models. However, for tasks such as factually consistent summarisation, existing metrics remain underdeveloped, limiting their effectiveness as signals for shaping model behaviour.While individual factuality metrics are unreliable, their combination can more effectively capture diverse factual errors. We leverage this insight to introduce an automated training pipeline that improves factual consistency in summaries by aggregating scores from different weak metrics. Our approach avoids the need for complex reward shaping by mapping scores to preferences and filtering out cases with high disagreement between metrics. For each source document, we generate lexically similar summary pairs by varying decoding strategies, enabling the model to learn from factual differences caused by subtle lexical differences. This approach constructs a high-quality preference dataset using only source documents.Experiments demonstrate consistent factuality gains across models, ranging from early encoder-decoder architectures to modern large language models, with smaller models reaching comparable factuality to larger ones.

摘要生成事实一致性偏好学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。