通过拼接不同文档上下文,提升小模型的主题可控摘要能力
Mix and Match: Context Pairing for Scalable Topic-Controlled Educational Summarisation
- 用不同文档的上下文配对生成对比训练样本
- 增广规模越大,摘要胜率和语义对齐度越高
- 小模型性能逼近大模型,适合资源受限场景
主题可控摘要允许用户生成聚焦于源文档特定方面的摘要。本文研究了一种针对小型语言模型(sLMs)进行主题可控摘要训练的数据增强策略。我们提出一种成对数据增强方法,将来自不同文档的上下文组合,生成对比性训练样本,使模型更有效地学习主题与摘要之间的关系。基于包含维基百科衍生主题的SciTLDR数据集,系统评估了增广规模对模型性能的影响。结果表明,随着增广规模增加,模型的胜率和语义对齐度持续提升,而真实训练数据量保持不变。因此,采用该增广方法训练的T5-base模型在参数量和真实训练样本数显著更少的情况下,仍达到与更大模型相当的性能。
原文摘要 · Abstract (English)
Topic-controlled summarisation enables users to generate summaries focused on specific aspects of source documents. This paper investigates a data augmentation strategy for training small language models (sLMs) to perform topic-controlled summarisation. We propose a pairwise data augmentation method that combines contexts from different documents to create contrastive training examples, enabling models to learn the relationship between topics and summaries more effectively. Using the SciTLDR dataset enriched with Wikipedia-derived topics, we systematically evaluate how augmentation scale affects model performance. Results show consistent improvements in win rate and semantic alignment as the augmentation scale increases, while the amount of real training data remains fixed. Consequently, a T5-base model trained with our augmentation approach achieves competitive performance relative to larger models, despite using significantly fewer parameters and substantially fewer real training examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。