arXiv:2604.25130cs.CL2026-04中稿 · vited paper for pu…被引 1

用问答反馈提升长文摘要评估与优化效果

LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization

论文配图:LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
图 1 · 摘自论文原文
  • 通过问答对衡量摘要可回答性与事实一致性
  • 在7个基准上比传统指标更贴近人工判断
  • 反馈可直接指导摘要自修正,无需重新训练

长文档摘要的评估仍是研究瓶颈。现有指标与人工判断相关性弱,仅给出笼统分数,无法解释缺陷或指导改进,难以支持需要可验证准确性的实际应用。我们提出LongSumEval,一个将评估与生成统一的框架,通过结构化问答反馈实现质量评估。该框架将摘要质量定义为问答对的答案可得性与事实一致性,生成可解释的评分和可操作的反馈,识别覆盖缺失与事实错误。元评估显示,我们的问答评估模块在七个基准上显著优于现有指标。结构化反馈可实现无重训练的自优化,大幅提升摘要质量。本工作证明评估反馈可作为生成的可执行指令,建立了一种可推广的评估-改进对齐范式,对需可验证准确性与透明质量控制的可控文本生成具有直接意义。所有代码与数据集将在GitHub开源。

原文摘要 · Abstract (English)

Evaluating long document summaries remains the primary bottleneck in summarization research. Existing metrics correlate weakly with human judgments and produce aggregate scores without explaining deficiencies or guiding improvement, preventing effective refinement in applications requiring verifiable accuracy. We introduce LongSumEval, a unified framework bridging evaluation and generation through structured question-answering feedback. The framework operationalizes summary quality as answerability and factual alignment of question-answer pairs, generating interpretable scores and actionable feedback that identifies coverage gaps and factual inconsistencies. This resolves the misalignment where evaluation operates independently of generation objectives. Meta-evaluation of our QA-based evaluation module across seven benchmarks demonstrates substantially stronger agreement with human judgments compared to established metrics. Structured feedback enables significant quality improvements through self-refinement without retraining. By demonstrating that evaluation feedback can serve as executable instructions for generation, this work establishes a generalizable paradigm for aligning assessment with improvement, with direct implications for controllable text generation requiring verifiable accuracy and transparent quality control. All code and datasets will be released in GitHub for reproducibility.

摘要评估问答系统自修正可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。