开源框架揭示摘要评估指标的可复现性问题
AllSummedUp: un framework open-source pour comparer les metriques d'evaluation de resume
- 构建统一开源框架,公平对比六种摘要评估指标
- 发现人类判断对齐度高的指标计算成本高且结果不稳定
- 适合关注评估可靠性的自然语言处理研究者使用
本文研究自动文本摘要评估中的可复现性挑战。在六个代表性评估指标(从传统的ROUGE到最新的基于大模型的方法G-Eval、SEval-Ex)上开展实验,发现文献中报告的性能与本研究实验设置下的结果存在显著差异。为此提出一个统一的开源框架,应用于SummEval数据集,旨在支持评估指标的公平透明比较。结果显示,与人类判断对齐度最高的指标往往计算开销大且跨运行稳定性差。研究还指出依赖大模型进行评估带来的随机性、技术依赖性和可复现性限制等问题,呼吁建立更稳健的评估协议,包括详尽文档和方法标准化,以提升自动摘要评估的可靠性。
原文摘要 · Abstract (English)
This paper investigates reproducibility challenges in automatic text summarization evaluation. Based on experiments conducted across six representative metrics ranging from classical approaches like ROUGE to recent LLM-based methods (G-Eval, SEval-Ex), we highlight significant discrepancies between reported performances in the literature and those observed in our experimental setting. We introduce a unified, open-source framework, applied to the SummEval dataset and designed to support fair and transparent comparison of evaluation metrics. Our results reveal a structural trade-off: metrics with the highest alignment with human judgments tend to be computationally intensive and less stable across runs. Beyond comparative analysis, this study highlights key concerns about relying on LLMs for evaluation, stressing their randomness, technical dependencies, and limited reproducibility. We advocate for more robust evaluation protocols including exhaustive documentation and methodological standardization to ensure greater reliability in automatic summarization assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。