用原子语句分解评估摘要质量,既准又可解释。
SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation
- 将摘要评价拆解为基本语句,通过语句匹配实现精准评估
- 在SummEval上一致性相关性达0.580,优于GPT-4的0.521
- 能生成评估依据,适合需要可解释性的研究与应用
文本摘要质量评估仍是自然语言处理中的关键挑战。现有方法在性能与可解释性之间存在权衡。本文提出SEval-Ex,一种基于语句级的可解释摘要评估框架,通过将评估任务分解为原子语句,实现高性能与高可解释性的统一。该框架采用两阶段流程:首先利用大语言模型从原文和摘要中提取原子语句,然后进行语句间匹配。与仅提供摘要级评分的现有方法不同,SEval-Ex通过语句级对齐生成详细的评估证据。在SummEval基准上的实验表明,SEval-Ex在一致性方面达到0.580的相关性,显著优于基于GPT-4的评估器(0.521),同时保持了良好的可解释性。最终,该框架对幻觉现象表现出强鲁棒性。
原文摘要 · Abstract (English)
Evaluating text summarization quality remains a critical challenge in Natural Language Processing. Current approaches face a trade-off between performance and interpretability. We present SEval-Ex, a framework that bridges this gap by decomposing summarization evaluation into atomic statements, enabling both high performance and explainability. SEval-Ex employs a two-stage pipeline: first extracting atomic statements from text source and summary using LLM, then a matching between generated statements. Unlike existing approaches that provide only summary-level scores, our method generates detailed evidence for its decisions through statement-level alignments. Experiments on the SummEval benchmark demonstrate that SEval-Ex achieves state-of-the-art performance with 0.580 correlation on consistency with human consistency judgments, surpassing GPT-4 based evaluators (0.521) while maintaining interpretability. Finally, our framework shows robustness against hallucination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。