用无标注对话数据持续预训练,提升大模型在电话摘要中的表现
DACP: Domain-Adaptive Continual Pre-Training of Large Language Models for Phone Conversation Summarization
- 通过持续预训练让大模型适应特定领域对话
- 在内外部测试集上均显著提升摘要质量
- 适合工业界缺乏标注数据的场景
大语言模型在文本摘要任务中表现优异,但在与原始预训练分布不同的专业领域中性能下降。微调虽可提升效果,但依赖昂贵且稀缺的高质量标注数据。本文探索了持续预训练这一可扩展的自监督方法,用于适应大模型在嘈杂真实对话转录数据上的摘要任务。我们利用大规模无标注业务对话数据进行实验,验证持续预训练是否能增强模型在对话摘要中的能力。结果表明,该方法在域内和域外摘要基准上均带来显著提升,同时保持强泛化性和鲁棒性。我们还分析了数据选择策略,为面向摘要的工业应用提供实用指导。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved impressive performance in text summarization, yet their performance often falls short when applied to specialized domains that differ from their original pre-training distribution. While fine-tuning can improve summarization quality, it typically relies on costly and scarce high-quality labeled data. In this work, we explore continual pre-training as a scalable, self-supervised approach to adapt LLMs for downstream summarization tasks, particularly in the context of noisy real-world conversation transcripts. We conduct extensive experiments using large-scale, unlabeled business conversation data to investigate whether continual pre-training enhances model capabilities in conversational summarization. Our results demonstrate that continual pre-training yields substantial gains in both in-domain and out-of-domain summarization benchmarks, while maintaining strong generalization and robustness. We also analyze the effects of data selection strategies, providing practical guidelines for applying continual pre-training in summarization-focused industrial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。