arXiv:2608.30609cs.CLcs.AI2026-08中稿 · EMNLP

用百万篇瑞典新闻微调大模型,提升新闻生成质量。

Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training

  • 在百万级瑞典新闻数据上持续预训练模型
  • 结合经验回放可缓解遗忘,显著提升生成与事实准确性
  • 专用评测集比现有基准更有效,尤其对低秩微调模型

大型语言模型在通用任务上表现优异,但在特定领域如瑞典新闻业的应用仍受限。本文通过在数百万篇新闻文章构建的高质量数据集上进行持续预训练,探索模型在瑞典新闻领域的适应性。为评估效果,我们构建了一个涵盖六种编辑任务的新领域专用基准。实验对比了两种规模模型的全量与参数高效微调,并发现持续预训练仅在结合经验回放时才能有效提升性能;模型在生成质量和事实知识方面均有改善,但判别类任务无明显进步。进一步研究显示,无需训练的指令引导方法仅在低秩适配(LoRA)训练的模型上带来增益。关键发现是:专用评测集能更准确捕捉领域内性能提升,而现有瑞典语基准则无法反映真实改进。

原文摘要 · Abstract (English)

Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to address this limitation is to specialise existing models through additional training on target-domain corpora. In this work, we investigate such continued pre-training for adapting large language models to Swedish journalism, using a high-quality dataset that we curate from millions of news articles. To evaluate the adaptation efficacy, we also construct a novel domain-specific benchmark that covers six editorial tasks. Through full and parameter-efficient fine-tuning across two model sizes, we find that continued pre-training yields benefits in the target domain, but only when paired with experience replay to mitigate forgetting. We observe consistent enhancements in the models' generation quality and factual knowledge, but not their proficiency in discriminative tasks. Exploring a training-free method to facilitate instruction following, we see further improvements, but exclusively for models trained with low-rank adaptation. Crucially, we demonstrate the importance of targeted evaluation in the adaptation process, as an existing Swedish benchmark largely fails to capture the models' in-domain performance gains.

语言模型新闻生成持续预训练瑞典语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。