arXiv:2601.16018cs.CL2026-01被引 4

专为土耳其法律领域打造的自研大模型,训练高效且性能领先。

Mecellem Models: Turkish Models Trained from Scratch and Continually Pre-trained for the Legal Domain

  • 从1127亿词元语料自研预训练编码器,单阶段完成,效率更高。
  • 小模型(155M)检索性能媲美大模型(307M-567M),92.36%生产效率领先。
  • 基于课程学习的持续预训练,使法律文本困惑度降低36.2%,适合法律AI研究者。

本文提出Mecellem模型框架,通过领域适配策略构建面向土耳其法律领域的专用语言模型。贡献有二:(1) 从头预训练编码器模型:基于ModernBERT的双向编码器,在1127亿词元的土耳其主导语料上训练。采用检查点选择策略,发现最优检查点在预训练损失达最低前即已取得最佳下游检索表现。小型模型(155M参数)在土耳其检索排行榜中位列前三,性能可比大型参考模型(307M-567M参数)。本方法生产效率达92.36%,优于SOTA模型(embeddinggemma-300m: 100.00%,BAAI/bge-m3: 99.54%,newmindai/bge-m3-stsb: 94.38%),排名第四,且所需算力更少。现有SOTA模型依赖多阶段、高成本训练流程,而本方法采用单阶段预训练+高效后训练,更具成本优势;(2) 带持续预训练(CPT)的解码器模型:使用Qwen3-1.7B与Qwen3-4B模型,通过受控课程学习适配至土耳其法律领域。四阶段CPT结合最优样本比例,实现从通用语言知识到专业法律术语与长上下文推理的渐进过渡。该方法在土耳其法律文本上实现36.2%的困惑度下降,验证了领域适配的有效性。

原文摘要 · Abstract (English)

This paper presents Mecellem models, a framework for developing specialized language models for the Turkish legal domain through domain adaptation strategies. We make two contributions: (1)Encoder Model Pre-trained from Scratch: ModernBERT-based bidirectional encoders pre-trained on a Turkish-dominant corpus of 112.7 billion tokens. We implement a checkpoint selection strategy that evaluates downstream retrieval performance throughout training, revealing that optimal checkpoints achieve best retrieval scores before pre-training loss reaches its minimum. Our encoder models achieve top-3 rankings on the Turkish retrieval leaderboard, with smaller models (155M parameters) achieving comparable performance to larger reference models (307M-567M parameters). Our approach achieves 92.36% production efficiency compared to state-of-the-art models (embeddinggemma-300m: 100.00%, BAAI/bge-m3: 99.54%, newmindai/bge-m3-stsb: 94.38%), ranking fourth overall despite requiring less computational resources. SOTA models rely on multi-stage, computationally intensive training pipelines, making our single-stage pre-training followed by efficient post-training approach a cost-effective alternative; (2)Decoder Model with Continual Pre-training (CPT): Qwen3-1.7B and Qwen3-4B models adapted to Turkish legal domain through controlled curriculum learning. Four-phase CPT with optimal sample ratios enables gradual transition from general language knowledge to specialized legal terminology and long-context reasoning. This approach achieves 36.2% perplexity reduction on Turkish legal text, demonstrating domain adaptation gains.

法律AI自研模型持续预训练土耳其

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。