用分层数据金字塔提升模型摘要与人类偏好一致性
AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference
- 构建包含三类摘要的数据金字塔,融合提取、生成与人工标注数据
- 经高斯重采样后,模型在CNN/DailyMail上超越1750亿参数GPT-3
- 适合追求高质量摘要、注重人类评价的NLP研究者使用
文本摘要任务通常使用预训练语言模型(PLMs)适配多样标准数据集。尽管这些模型在自动评估中表现优异,但在人类评估中常表现不佳,表明生成摘要与人类摘要偏好存在偏差。这可能源于微调数据质量低以及真实人类偏好标注数据稀缺。为此,我们提出新型人类摘要偏好对齐框架AlignSum。该框架包含三部分:首先,构建包含抽取式、抽象式和人工标注摘要的数据金字塔;其次,采用高斯重采样剔除极端长度摘要;最后,在高斯重采样后实施两阶段分层微调。我们在人工标注的CNN/DailyMail和BBC XSum数据集上应用AlignSum。实验表明,经AlignSum微调的BART-Large模型在自动与人工评估中均超越1750亿参数的GPT-3,证明该方法显著提升语言模型与人类摘要偏好的对齐度。
原文摘要 · Abstract (English)
Text summarization tasks commonly employ Pre-trained Language Models (PLMs) to fit diverse standard datasets. While these PLMs excel in automatic evaluations, they frequently underperform in human evaluations, indicating a deviation between their generated summaries and human summarization preferences. This discrepancy is likely due to the low quality of fine-tuning datasets and the limited availability of high-quality human-annotated data that reflect true human preference. To address this challenge, we introduce a novel human summarization preference alignment framework AlignSum. This framework consists of three parts: Firstly, we construct a Data Pymarid with extractive, abstractive, and human-annotated summary data. Secondly, we conduct the Gaussian Resampling to remove summaries with extreme lengths. Finally, we implement the two-stage hierarchical fine-tuning with Data Pymarid after Gaussian Resampling. We apply AlignSum to PLMs on the human-annotated CNN/DailyMail and BBC XSum datasets. Experiments show that with AlignSum, PLMs like BART-Large surpass 175B GPT-3 in both automatic and human evaluations. This demonstrates that AlignSum significantly enhances the alignment of language models with human summarization preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。