arXiv:2502.09642cs.CLcs.AI2025-02被引 13

为印度10亿人打造的多语言大模型,解决语言数据稀缺问题。

Krutrim LLM: Multilingual Foundational Model for over a Billion People

  • 基于2万亿令牌训练,整合最大已知印地语数据集。
  • 在16项任务中10项超越或持平LLAMA-2,平均得分0.57。
  • 支持实时搜索,提升对话AI准确性,适合多语言用户。

印度社会语言多样、数据获取难、口语传统丰富,现有基础模型多以英语训练,难以服务其18%全球人口占比。印地语仅占Common Crawl语料库1%,造成显著语言偏见。数千种方言与代码混用导致训练数据稀疏。我们提出Krutrim LLM,一个训练量达2万亿令牌的多语言模型,涵盖最大已知印地语数据集,缓解数据短缺并实现各方言均衡表现。在印地语基准测试中,其性能优于或匹配当前最优模型,同时保持与英文模型相当的表现。尽管训练浮点运算量显著更少,其在16项任务中的10项表现优于或持平LLAMA-2,平均得分为0.57(对比0.55)。通过集成实时搜索,提升对话式AI的准确性,惠及超10亿用户。该模型通过针对性设计克服数据不平衡,推动构建更具伦理性和全球代表性的AI系统。

原文摘要 · Abstract (English)

India is a diverse society with unique challenges in developing AI systems, including linguistic diversity, oral traditions, data accessibility, and scalability. Existing foundation models are primarily trained on English, limiting their effectiveness for India's population. Indic languages comprise only 1 percent of Common Crawl corpora despite India representing 18 percent of the global population, leading to linguistic biases. Thousands of regional languages, dialects, and code mixing create additional representation challenges due to sparse training data. We introduce Krutrim LLM, a 2 trillion token multilingual model designed for India's linguistic landscape. It incorporates the largest known Indic dataset, mitigating data scarcity and ensuring balanced performance across dialects. Krutrim outperforms or matches state-of-the-art models on Indic benchmarks while maintaining competitive English performance. Despite being significantly smaller in training flops, Krutrim LLM matches or exceeds models like LLAMA-2 on 10 out of 16 tasks, with an average score of 0.57 versus 0.55. This evidences Krutrim's flexible multilingual fluency across diverse linguistic contexts. Krutrim is integrated with real-time search to improve factual accuracy in conversational AI applications. This enhances accessibility for over 1 billion users worldwide. Through intentional design choices addressing data imbalances, Krutrim LLM signifies meaningful progress in building ethical, globally representative AI models.

多语言模型印地语数据平衡对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。