更新的法语模型CamemBERT 2.0更懂现代法语,性能显著提升。
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
- 基于DeBERTaV3和RoBERTa架构,采用新训练目标提升理解力
- 在医疗等专业任务中表现优于旧版,大幅减少过时数据影响
- 适合需要最新法语理解能力的研究与工业应用
法语模型如CamemBERT已广泛应用于自然语言处理任务,月下载量超400万次。然而,由于训练数据过时导致的语义漂移问题,模型在面对新话题和术语时性能下降。本文推出两个新版本的CamemBERT base模型:CamemBERTav2(基于DeBERTaV3,使用替换词检测目标)和CamemBERTv2(基于RoBERTa,使用掩码语言建模)。两者均在更大、更近期的数据集上训练,支持更长上下文,并采用更新的分词器。在通用及医疗等特定领域任务中评估显示,新模型性能远超前代,具备更强适应性与实用性。所有新模型及中间检查点均已开源至Huggingface。
原文摘要 · Abstract (English)
French language models, such as CamemBERT, have been widely adopted across industries for natural language processing (NLP) tasks, with models like CamemBERT seeing over 4 million downloads per month. However, these models face challenges due to temporal concept drift, where outdated training data leads to a decline in performance, especially when encountering new topics and terminology. This issue emphasizes the need for updated models that reflect current linguistic trends. In this paper, we introduce two new versions of the CamemBERT base model-CamemBERTav2 and CamemBERTv2-designed to address these challenges. CamemBERTav2 is based on the DeBERTaV3 architecture and makes use of the Replaced Token Detection (RTD) objective for better contextual understanding, while CamemBERTv2 is built on RoBERTa, which uses the Masked Language Modeling (MLM) objective. Both models are trained on a significantly larger and more recent dataset with longer context length and an updated tokenizer that enhances tokenization performance for French. We evaluate the performance of these models on both general-domain NLP tasks and domain-specific applications, such as medical field tasks, demonstrating their versatility and effectiveness across a range of use cases. Our results show that these updated models vastly outperform their predecessors, making them valuable tools for modern NLP systems. All our new models, as well as intermediate checkpoints, are made openly available on Huggingface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。