arXiv:2602.21379cs.CLcs.AI2026-02被引 1

MrBERT通过多语言适配,实现小参数量下的多语种高精度与高效推理。

MrBERT: Modern Multilingual Encoders via Vocabulary, Domain, and Dimensional Adaptation

  • 基于ModernBERT架构,融合词汇、领域与维度自适应优化
  • 在加泰罗尼亚语、西班牙语任务上达顶尖表现,生物医学/法律领域稳健
  • 支持可变向量尺寸,显著降低推理与存储开销,适合生产部署

我们提出MrBERT,一个基于ModernBERT架构的150M-300M参数编码器家族,覆盖35种语言和代码。通过针对性适配,该模型在加泰罗尼亚语和西班牙语特定任务中达到当前最优性能,并在生物医学和法律等专业领域表现稳健。为缩小研究与生产差距,引入马特里什卡表示学习(MRL),支持灵活向量尺寸,大幅降低推理与存储成本。最终证明,现代编码器架构可通过优化实现本地语言优势与高价值领域专精的兼顾。完整模型家族已在Huggingface开源。

原文摘要 · Abstract (English)

We introduce MrBERT, a family of 150M-300M parameter encoders built on the ModernBERT architecture and pre-trained on 35 languages and code. Through targeted adaptation, this model family achieves state-of-the-art results on Catalan- and Spanish-specific tasks, while establishing robust performance across specialized biomedical and legal domains. To bridge the gap between research and production, we incorporate Matryoshka Representation Learning (MRL), enabling flexible vector sizing that significantly reduces inference and storage costs. Ultimately, the MrBERT family demonstrates that modern encoder architectures can be optimized for both localized linguistic excellence and efficient, high-stakes domain specialization. We open source the complete model family on Huggingface.

多语言编码器模型压缩开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。