构建医疗摘要新基准,精准追踪每句话的来源依据。
PCoA: A New Benchmark for Medical Aspect-Based Summarization With Phrase-Level Context Attribution
- 提出细粒度标注体系,关联摘要与原文中具体语句和短语。
- 实验证明,先定位相关句再生成摘要可提升整体质量。
- 适合医学文本生成、可解释性评估等研究者使用。
验证自动生成的摘要仍具挑战性,因有效验证需精确追溯到原始上下文,尤其在高风险医疗领域至关重要。为此,我们提出PCoA——一个专家标注的医疗领域基于方面摘要的基准数据集,支持短语级上下文归因。PCoA将每个摘要方面与其对应的上下文句子及其中贡献短语精确对齐。我们进一步设计了一种细粒度、解耦的评估框架,独立评估生成摘要、引用和贡献短语的质量。通过大量实验,我们验证了PCoA数据集的高质量与一致性,并在该任务上评测多个大语言模型。实验结果表明,PCoA为具有短语级上下文归因的系统摘要评估提供了可靠基准。此外,对比实验显示,在摘要生成前显式识别相关句子与贡献短语,能显著提升整体生成质量。数据与代码已公开于https://github.com/chubohao/PCoA。
原文摘要 · Abstract (English)
Verifying system-generated summaries remains challenging, as effective verification requires precise attribution to the source context, which is especially crucial in high-stakes medical domains. To address this challenge, we introduce PCoA, an expert-annotated benchmark for medical aspect-based summarization with phrase-level context attribution. PCoA aligns each aspect-based summary with its supporting contextual sentences and contributory phrases within them. We further propose a fine-grained, decoupled evaluation framework that independently assesses the quality of generated summaries, citations, and contributory phrases. Through extensive experiments, we validate the quality and consistency of the PCoA dataset and benchmark several large language models on the proposed task. Experimental results demonstrate that PCoA provides a reliable benchmark for evaluating system-generated summaries with phrase-level context attribution. Furthermore, comparative experiments show that explicitly identifying relevant sentences and contributory phrases before summarization can improve overall quality. The data and code are available at https://github.com/chubohao/PCoA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。