用少数据高效训练德语工业领域语言模型,省时又提效。
Efficient Domain-adaptive Continual Pretraining for the Process Industry in the German Language
- 结合上下文学习与最近邻技术,补足工业德语文本
- 性能比当前最优方法高7.87点,显存消耗降近四倍
- 适合算力有限的制造业企业快速落地NLP应用
领域自适应持续预训练(DAPT)是一种先进方法,可在缺乏标注任务数据时,通过继续在预训练任务(如掩码语言建模,MLM)上训练语言模型来提升性能。然而,MLM需要大量领域相关文本,这对非英语领域(如德语工业领域)而言难以获取。本文提出一种高效方法ICL-augmented pretraining(ICL-APT),利用上下文学习(ICL)和k近邻(kNN)技术,从外部来源生成并增强目标领域的文本数据,显著降低GPU使用时间。实验表明,ICL-APT最佳配置相比现有最先进方法提升28.7%(7.87分),且所需GPU计算时间几乎减少4倍,为计算资源受限的产业提供了一种低成本、高效的解决方案。该框架对其他低资源行业也具广泛适用性,推动NLP技术在生产环境中的实际部署。
原文摘要 · Abstract (English)
Domain-adaptive continual pretraining (DAPT) is a state-of-the-art technique that further trains a language model (LM) on its pretraining task, e.g., masked language modeling (MLM), when common domain adaptation via LM fine-tuning is not possible due to a lack of labeled task data. Although popular, MLM requires a significant corpus of domain-related data, which is difficult to obtain for specific domains in languages other than English, such as the process industry in the German language. This paper introduces an efficient approach called ICL-augmented pretraining or ICL-APT that leverages in-context learning (ICL) and k-nearest neighbors (kNN) to augment target data with domain-related and in-domain texts, significantly reducing GPU time while maintaining strong model performance. Our results show that the best configuration of ICL-APT performed better than the state-of-the-art DAPT by 28.7% (7.87 points) and requires almost 4 times less GPU-computing time, providing a cost-effective solution for industries with limited computational capacity. The findings highlight the broader applicability of this framework to other low-resource industries, making NLP-based solutions more accessible and feasible in production environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。