arXiv:2609.07798cs.CLcs.AI2026-09

用三类气候文本优化预训练模型,提升气候领域NLP性能

Climate-ModernBERT: Revisiting Corpus Composition for Domain-Adaptive Continued Pretraining

论文配图:Climate-ModernBERT: Revisiting Corpus Composition for Domain-Adaptive Continued Pretraining
图 1 · 摘自论文原文
  • 在学术、网络和合成气候文本上联合微调模型
  • 平均F1达76.3,较基线提升2.8点
  • 参数空间融合优于多源联合训练,适合气候领域研究

气候领域自然语言处理需处理科学文献、政策文件和合成报告等异构文本。然而,如何在持续预训练中有效整合多样领域语料仍缺乏探索。本文提出Climate-ModernBERT,基于ModernBERT-Base在三种气候语料(学术气候文本、过滤后的网络数据、合成气候文档)上进行持续预训练。系统比较了联合多源训练与独立微调后参数空间融合的策略。在九个气候NLP基准上,最优模型平均F1为76.3,较原始ModernBERT提升2.8点。结果表明,学术气候语料提供最强适应信号;参数空间融合优于联合训练,更好保留异构语料的互补信息。所有Climate-ModernBERT变体及训练检查点已开源,支持气候NLP与领域自适应预训练研究。

原文摘要 · Abstract (English)

Natural Language Processing (NLP) in the climate domain requires models to process heterogeneous text sources, including scientific literature, policy disclosures, and synthetic reports. However, how to effectively combine diverse domain corpora during continued pretraining (CPT) remains underexplored. We introduce Climate-ModernBERT, a family of climate-adapted encoder models obtained through continued pretraining of ModernBERT-Base on three climate corpora: academic climate text, climate-filtered web data, and synthetic climate documents. We systematically compare joint continued pretraining on corpus mixtures with parameter-space merging of independently specialized checkpoints. Across nine climate NLP benchmarks, our best model achieves 76.3 average F_1, improving significantly over a vanilla ModernBERT baseline by 2.8 points. Within the climate NLP setting, the results show that academic climate corpora provide the strongest adaptation signal among the evaluated sources, while parameter-space merging improves over joint multi-source training and better preserves complementary information from heterogeneous climate corpora. We release all Climate-ModernBERT variants and training checkpoints to support future research in climate NLP and domain-adaptive pretraining.

气候NLP持续预训练模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。