arXiv:2603.14563cs.CL2026-03

为17种印度语言打造儿童故事数据集,助力小模型训练

Multilingual TinyStories: A Synthetic Combinatorial Corpus of Indic Children's Stories for Training Small Language Models

  • 用混合方法合成故事:自研提示框架+大模型生成+翻译扩增
  • 产出13万+故事、超9300万词元,覆盖17种印度语言
  • 专为小模型设计,适合低资源语言研究与跨语言迁移

低资源语言的鲁棒语言模型发展常受限于高质量、连贯且领域适配的训练语料。本文提出多语言TinyStories数据集,一个大规模合成的儿童故事集合,涵盖17种印度语言。该语料专为小型语言模型(SLMs)的训练与评估设计,内容以简单叙事为主,严格使用母语文字。我们采用混合数据筛选流程,结合Sarvam-M语言模型与新颖的组合式提示工程框架实现本地化生成,并通过Google Translate API实现大规模跨语言扩展。经严格程序化过滤后,最终整理出132,942个故事及超过9390万词元的语料,为印地语系多语言建模与迁移学习提供基础资源。

原文摘要 · Abstract (English)

The development of robust language models for low-resource languages is frequently bottlenecked by the scarcity of high-quality, coherent, and domain-appropriate training corpora. In this paper, we introduce the Multilingual TinyStories dataset, a large-scale, synthetically generated collection of children's stories encompassing 17 Indian languages. Designed specifically for the training and evaluation of Small Language Models (SLMs), the corpus provides simple, narrative-driven text strictly localized to native scripts. We detail our hybrid curation pipeline, which leverages the Sarvam-M language model and a novel combinatorial prompt engineering framework for native generation, coupled with the Google Translate API for large-scale cross-lingual expansion. Through strict programmatic filtering, we compiled 132,942 stories and over 93.9 million tokens in our release, serving as a foundational resource for multilingual language modeling and transfer learning in the Indic linguistic sphere.

多语言小模型合成数据儿童故事

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。