开源100亿参数的印度语大模型,专为印度语生成优化
Llama-3-Nanda-10B-Chat: An Open Generative Large Language Model for Hindi
- 基于Llama-3架构,通过数据增强和双语训练提升印度语能力
- 100亿参数,性能超越同类开源印度语模型
- 适合研究印度语AI、本地化应用及多语言系统开发
针对中等资源语言的高质量大语言模型开发面临数据稀缺、模型适配与评估难题。本文推出基于Llama-3-8B的开源生成式大模型Llama-3-Nanda-10B-Chat(简称Nanda),专为印地语优化。通过扩展变换器结构进行连续预训练,并采用严格的文本数据清洗、增强与策略性双语训练,平衡印地语与英语语料以促进跨语言知识迁移。该模型共100亿参数,在同等规模开源印地语与多语言模型中表现领先,显著优于现有多数模型。论文深入探讨了训练策略、微调方法、安全对齐与评估指标,展示了其达到顶尖水平的关键路径。通过开源,旨在推动印地语大模型研究,并支持学术、产业及公共服务中的实际应用。
原文摘要 · Abstract (English)
Developing high-quality large language models (LLMs) for moderately resourced languages presents unique challenges in data availability, model adaptation, and evaluation. We introduce Llama-3-Nanda-10B-Chat, or Nanda for short, a state-of-the-art Hindi-centric instruction-tuned generative LLM, designed to push the boundaries of open-source Hindi language models. Built upon Llama-3-8B, Nanda incorporates continuous pre-training with expanded transformer blocks, leveraging the Llama Pro methodology. A key challenge was the limited availability of high-quality Hindi text data; we addressed this through rigorous data curation, augmentation, and strategic bilingual training, balancing Hindi and English corpora to optimize cross-linguistic knowledge transfer. With 10 billion parameters, Nanda stands among the top-performing open-source Hindi and multilingual models of similar scale, demonstrating significant advantages over many existing models. We provide an in-depth discussion of training strategies, fine-tuning techniques, safety alignment, and evaluation metrics, demonstrating how these approaches enabled Nanda to achieve state-of-the-art results. By open-sourcing Nanda, we aim to advance research in Hindi LLMs and support a wide range of real-world applications across academia, industry, and public services.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。