arXiv:2506.15925cs.CL2025-06ACL被引 2

用重排序提升无偏视角摘要质量,评测更准。

Reranking-based Generation for Unbiased Perspective Summarization

  • 用人类标注数据测试指标可靠性,发现语言模型指标更优。
  • 重排序方法显著提升摘要质量,合成数据微调效果更好。
  • 适合关注公平性与可信赖评估的NLP研究者。

在真实场景如政治视角摘要生成中,实现无偏总结仍是大语言模型(LLMs)的重要应用。然而,现有评估框架依赖传统指标衡量覆盖度和忠实性,未验证其适用性,且改进摘要生成方法仍处于初步阶段。本文通过(1)识别可靠的视角摘要质量评估指标,(2)探究基于LLM的方法超越零样本推理的潜力。我们构建了一个用于基准测试指标可靠性的测试集,结合人工标注,发现传统指标表现逊于基于语言模型的指标,后者表现出更强的评估能力。利用这些指标,我们证实重排序方法能取得优异结果,且通过合成数据与重排序标注进行偏好微调可进一步提升性能。研究旨在推动视角摘要方法的可靠评估与持续发展。

原文摘要 · Abstract (English)

Generating unbiased summaries in real-world settings such as political perspective summarization remains a crucial application of Large Language Models (LLMs). Yet, existing evaluation frameworks rely on traditional metrics for measuring key attributes such as coverage and faithfulness without verifying their applicability, and efforts to develop improved summarizers are still nascent. We address these gaps by (1) identifying reliable metrics for measuring perspective summary quality, and (2) investigating the efficacy of LLM-based methods beyond zero-shot inference. Namely, we build a test set for benchmarking metric reliability using human annotations and show that traditional metrics underperform compared to language model-based metrics, which prove to be strong evaluators. Using these metrics, we show that reranking-based methods yield strong results, and preference tuning with synthetically generated and reranking-labeled data further boosts performance. Our findings aim to contribute to the reliable evaluation and development of perspective summarization methods.

摘要生成无偏性评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。