arXiv:2603.03508cs.CLcs.AI2026-03

0.6B参数的独立训练模型,让印地语语言模型性能超越大模型。

Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi

  • 从零开始构建印地语专用模型,全程透明可复现。
  • 在多个评测中优于同规模多语言模型,如Qwen2.5-0.5B。
  • 适合资源有限但需高质量印地语模型的研究与应用。

大型多语言基础模型的主导地位加剧了自然语言处理中的语言不平等,常导致低资源语言被忽视。本文提出LilMoo,一个0.6亿参数的印地语语言模型,完全从头训练以弥补这一差距。不同于依赖不透明多语言预训练的现有印地语模型,LilMoo采用全透明、可复现的训练流程,专为计算资源受限环境优化。我们构建了一个高质量印地语语料库GigaLekh,通过启发式与基于大模型评判(LLM-as-a-judge)的方法筛选,并辅以精选英文数据进行双语增强。利用该数据集,探索小规模语言模型的多种训练方案。在全面评估中,LilMoo持续优于同等规模的多语言基线模型,如Qwen2.5-0.5B和Qwen3-0.6B,证明精心设计的语言特定预训练可在子十亿参数范围内媲美大型多语言模型。

原文摘要 · Abstract (English)

The dominance of large multilingual foundation models has widened linguistic inequalities in Natural Language Processing (NLP), often leaving low-resource languages underrepresented. This paper introduces LilMoo, a 0.6-billion-parameter Hindi language model trained entirely from scratch to address this gap. Unlike prior Hindi models that rely on continual pretraining from opaque multilingual foundations, LilMoo is developed through a fully transparent and reproducible pipeline optimized for limited compute environments. We construct a high-quality Hindi corpus (GigaLekh) filtered through both heuristic and learned (LLM-as-a-judge) methods, complemented by bilingual augmentation with curated English data. Using this dataset, we explore various training recipes for small-scale language models. Across comprehensive evaluation suites, LilMoo consistently outperforms comparably sized multilingual baselines such as Qwen2.5-0.5B and Qwen3-0.6B, demonstrating that well-designed language-specific pretraining can rival large multilingual models at the sub-billion-parameter range.

印地语小模型语言公平自训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。