用小模型对比印地语等区域语言的词法与学习效果
Regional Tiny Stories: Using Small Models to Compare Language Learning and Tokenizer Performance
- 用大模型生成合成数据,扩展小模型在印地语等语言上的训练
- 小模型仅需百万参数即可生成连贯文本,且印地语表现优于马拉地语和孟加拉语
- 专用分词器比通用分词器更适合印度语言,适合资源匮乏语言研究者
小语言模型(SLMs)为特定领域提供了高效替代方案。2023年TinyStories研究构建了英文数据集,使100万至1000万参数的小模型可生成连贯输出。本研究将该框架拓展至印度语言,翻译原始数据集并利用大模型生成合成数据,聚焦印地语、马拉地语和孟加拉语。结果表明,小模型以显著更少参数即可有效处理区域语言,为基于推理的分词策略与语言复杂性评估提供新范式。分析显示,针对语言特性的分词器优于通用分词器;信息论与形态学分析揭示印地语模型表现优于马拉地语和孟加拉语的原因。此外,合成数据在训练小模型时优于直接翻译内容。相关性分析揭示跨语言模式及创造力、语法精确性与叙事完整性间的语言特异性关系。研究推动小模型在低资源语言中的应用,并深化对神经语言发展的理论理解。
原文摘要 · Abstract (English)
Small Language Models (SLMs) offer efficient alternatives to LLMs for specific domains. The 2023 TinyStories study developed an English dataset that allows SLMs with 1 to 10 million parameters to produce coherent outputs. Our research expands this framework by translating the original dataset into Indian languages and creating synthetic data using LLMs. We focus on Hindi, Marathi, and Bengali, evaluating SLMs for regional language processing and understanding linguistic complexity. We show that SLMs efficiently process regional languages with significantly fewer parameters than LLMs, providing a complementary framework for ``inference based evaluation" of tokenization strategies and linguistic complexity. Our analysis shows that language-specific tokenizers outperform general-purpose ones for Indian languages. Empirical validations, supported by information-theoretic and morphological analyses, provides fundamental understanding behind the better performance of Hindi models over Marathi and Bengali. Additionally, we show that synthetic datasets outperform translated content for training SLMs. Correlation analyses reveal cross-linguistic patterns and language-specific relationships between creativity, grammatical precision, and narrative completeness. These findings advance both the practical application of SLMs to underserved languages and our theoretical understanding of neural language development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。