无需参考摘要即可准确评估摘要质量,提升模型评分可靠性。
Calibrating Model-Based Evaluation Metrics for Summarization

- 提出新框架,不依赖参考摘要或人工标注生成评分
- 使用组等距回归分箱法校准预测值,显著提升准确性
- 适用于摘要、问答等任务,尤其适合缺乏参考文本场景
近期摘要评估方法多基于大语言模型,评估完整性、简洁性与忠实度等维度。然而这些方法常需大量计算资源,且预测分数常出现偏差,影响可信度。此外,评估单文档多个摘要的平均质量通常需要多个参考摘要。本文提出一种通用框架,可在无需参考摘要、人工标注或昂贵模型的情况下,生成个体与平均代理评分。同时提出组等距回归分箱(GIRB)校准方法,使原始预测更贴近真实评价指标。尽管聚焦于连续值任务如摘要生成,该方法亦适用于离散值任务如问答。在七个数据集上的实验表明,本方法持续优于现有基线。
原文摘要 · Abstract (English)
Recent advances in summary evaluation are based on model-based metrics to assess quality dimensions, such as completeness, conciseness, and faithfulness. However, these methods often require large language models, and predicted scores are frequently miscalibrated, limiting their reliability. Moreover, evaluating the average quality across different summaries for a single document typically requires access to multiple reference summaries. Here, we propose a general framework that generates individual and average proxy scores without relying on reference summaries, human annotations, or expensive model-based metrics. We also propose group isotonic regression binning (GIRB), a calibration method that adjusts the raw predictions to better align with ground-truth evaluation metrics. While we focus on continuous-value scenarios, such as summarization, the method is applicable to discrete-value tasks, such as question answering. Experiments on seven datasets demonstrate that our approach consistently outperforms existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。