用合成数据和人工验证,让30亿参数模型学会尼泊尔/印度的土著语言Tharu。
TharuChat: Bootstrapping Large Language Models for a Low-Resource Language via Synthetic Data and Human Validation
- 用谷歌Gemini生成带方言特色的合成文本,再由母语者校验。
- 数据量提升至原来的4倍,困惑度从6.42降到2.88,模型更准确。
- 为濒危语言提供低成本保护方案,适合资源有限的研究者参考。
大型语言模型的快速普及加剧了全球南方原住民语言被排除在人工智能革命之外的数字鸿沟。以尼泊尔和印度特里地区约170万人口使用的土著语言Tharu为例,尽管有丰富的口述传统,但其面临严重数据匮乏与语言碎片化问题,导致主流多语言模型常出现幻觉或默认使用高资源语言如印地语和尼泊尔语。本文提出Tharu-LLaMA(3B),一种专为该语言设计的指令跟随模型。我们构建了TharuChat数据集,通过提示工程驱动的Gemini模型生成合成数据,并结合拉纳Tharu语法与民俗内容。该数据集主要基于拉纳Tharu(约70%),同时融合达那古拉和科奇拉方言元素。我们对数据局限性进行了透明分析,包括方言混杂与残留阿瓦德语/印地语影响。实证消融实验表明,即使存在这些缺陷,小规模合成数据仍具高度有效性:数据量从25%增至100%,困惑度从6.42线性下降至2.88。所建模型为利用生成式AI保护喜马拉雅低资源语言提供了概念验证,且可在消费级硬件上实现。
原文摘要 · Abstract (English)
The rapid proliferation of Large Language Models (LLMs) has created a profound digital divide, effectively excluding indigenous languages of the Global South from the AI revolution. The Tharu language, an Indo-Aryan vernacular spoken by approximately 1.7 million people across the Terai belt of Nepal and India, exemplifies this crisis. Despite a rich oral tradition, Tharu suffers from severe data scarcity and linguistic fragmentation, causing state-of-the-art multilingual models to routinely "hallucinate" or default to dominant high-resource neighbors like Hindi and Nepali due to contamination in pre-training corpora. This paper presents Tharu-LLaMA (3B), a specialized instruction-following model designed to address this exclusion. We introduce TharuChat, a novel dataset constructed via a LLM-to-Human bootstrapping pipeline. We utilized prompt-engineered Gemini models, fed with Rana Tharu grammar and folklore, to synthesize training data. Unlike curated gold-standard corpora, TharuChat reflects the noisy, heterogeneous linguistic reality of the region: it is predominantly anchored in Rana Tharu (~70%) while integrating elements of Dangaura and Kochila dialects. We provide a transparent analysis of the dataset's limitations, including dialectal code-mixing and residual Awadhi/Hindi influence. Through a rigorous empirical ablation study, we demonstrate that despite these imperfections, small-scale synthetic data is highly effective, increasing the dataset volume from 25% to 100% results in a linear reduction in perplexity from 6.42 to 2.88. The resulting model serves as a proof-of-concept for the preservation of under-resourced Himalayan languages via generative AI, achievable on consumer-grade hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。