筛选高质量摘要数据,提升科学文献摘要生成效果。
Less is More: Quality-Aware Training Data Selection for Scientific Summarization

- 用来源和模型双重指标评估作者摘要质量
- 选优质数据训练,相同规模下效果优于随机采样
- 适合研究长文本摘要与数据质量的学者
科学长文档摘要数据集通常将作者撰写的摘要作为标准参考,但其质量和与原文的一致性参差不齐。同时,现有公开科学摘要数据集在规模和结构上仍难以满足现代长上下文模型的需求。本文通过a) 构建并发布目前最大的生物医学与生命科学长文档摘要数据集,包含188万篇PMC文章;b) 利用基于来源和模型的度量方法分析作者摘要的质量。结果表明,作者摘要与全文对齐程度差异显著,这些质量信号可指导训练数据选择。在相同训练规模下,使用高质量子集训练的表现优于随机采样,且在事实准确性指标上可达到甚至超过更大规模的随机子集。研究提示参考摘要质量是科学摘要的关键因素,质量感知的数据选择能有效提升训练效率。
原文摘要 · Abstract (English)
Scientific long-document summarization datasets commonly treat author-written abstracts as gold reference summaries, although their quality and alignment with the source article vary. At the same time, publicly available scientific summarization datasets remain limited in scale and structure for modern long-context models. In this work, we address both challenges by a) constructing and releasing one of the largest biomedical and life science datasets for long-document summarization, containing 1.88 million PMC articles, and b) analyzing the reference quality of author-written abstracts with source-grounded and model-based metrics. We show that author-written abstracts vary in their alignment with the full article and that these quality signals can guide training-data selection. Training on selected high-quality subsets outperforms random sampling at matched training sizes and can match or exceed larger random subsets on factuality-oriented metrics. Our findings suggest that reference quality is an important factor in scientific summarization and that quality-aware data selection can improve training efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。