arXiv:2507.13390cs.CLcs.LG2025-07被引 4

2.9B参数模型专注印度多语言,让本地语言在大模型中真正被看见。

PARAM-1 BharatGen 2.9B Model

  • 从零训练的解码器模型,专为印地语和英语双语设计
  • 25%数据来自印地语,支持代码混用与方言适应性任务
  • 适合关注印度本土化应用的研究者与开发者

大型语言模型虽具强大推理能力,但其发展仍以英语为中心,导致印度等多语言地区长期被边缘化。印度拥有超过20种官方语言和100多种方言,普遍存在语码转换与双言现象。本文提出PARAM-1,一个2.9B参数的纯文本解码器模型,从头训练,专注于印度语言多样性。模型使用仅包含印地语和英语的双语语料库,强调事实丰富、高质量内容。遵循三大原则:印地语类语言占25%语料;采用适配印度形态结构的SentencePiece分词器保障分词公平性;在IndicQA、混合语言推理与社会语言鲁棒性任务上建立文化对齐评估基准。通过将多样性嵌入预训练阶段而非后期对齐,PARAM-1提供了一种以设计为核心的公平基础建模范式。实验表明,它既是通用能力强的模型,也是面向印度场景的可靠基线。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as powerful general-purpose reasoning systems, yet their development remains dominated by English-centric data, architectures, and optimization paradigms. This exclusionary design results in structural under-representation of linguistically diverse regions such as India, where over 20 official languages and 100+ dialects coexist alongside phenomena like code-switching and diglossia. We introduce PARAM-1, a 2.9B parameter decoder-only, text-only language model trained from scratch with an explicit architectural and linguistic focus on Indian diversity. PARAM-1 is trained on a bilingual dataset consisting of only Hindi and English, constructed with a strong focus on fact-rich, high-quality content. It is guided by three core principles: equitable representation of Indic languages through a 25% corpus allocation; tokenization fairness via a SentencePiece tokenizer adapted to Indian morphological structures; and culturally aligned evaluation benchmarks across IndicQA, code-mixed reasoning, and socio-linguistic robustness tasks. By embedding diversity at the pretraining level-rather than deferring it to post-hoc alignment-PARAM-1 offers a design-first blueprint for equitable foundation modeling. Our results demonstrate that it serves as both a competent general-purpose model and a robust baseline for India-centric applications.

多语言模型印度语公平性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。