arXiv:2512.03903cs.CLcs.AI2025-12

为巴斯克语构建多样化语料库,提升语言模型的泛化能力

BERnaT: Basque Encoders for Representing Natural Textual Diversity

  • 基于标准、社交媒体和历史文本构建巴斯克语语料库
  • 混合训练模型在标准与非标准任务上均表现更优
  • 适合关注语言多样性与模型公平性的研究者

语言模型依赖大规模文本语料,但质量过滤常无意排除非标准语言变体,降低模型鲁棒性并强化表征偏见。本文主张语言模型应捕捉语言变异的全谱(方言、历史、非正式等),而非仅依赖标准化文本。聚焦巴斯克语,我们整合标准、社交媒体及历史文本构建新语料库,并预训练BERnaT系列编码器模型,包含三种配置:标准型、多样型与混合型。提出评估框架,将自然语言理解任务分为标准与多样子集,以衡量语言泛化能力。结果表明,同时使用标准与多样化数据训练的模型,在所有任务类型中均优于仅用标准数据训练的模型,且不牺牲标准基准性能。研究凸显了语言多样性在构建包容性、泛化能力强的语言模型中的重要性。

原文摘要 · Abstract (English)

Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally exclude non-standard linguistic varieties, reduce model robustness and reinforce representational biases. In this paper, we argue that language models should aim to capture the full spectrum of language variation (dialectal, historical, informal, etc.) rather than relying solely on standardized text. Focusing on the Basque language, we construct new corpora combining standard, social media, and historical sources, and pre-train the BERnaT family of encoder-only models in three configurations: standard, diverse, and combined. We further propose an evaluation framework that separates Natural Language Understanding (NLU) tasks into standard and diverse subsets to assess linguistic generalization. Results show that models trained on both standard and diverse data consistently outperform those trained on standard corpora, improving performance across all task types without compromising standard benchmark accuracy. These findings highlight the importance of linguistic diversity in building inclusive, generalizable language models.

语言模型多样性巴斯克语编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。