arXiv:2511.10338cs.CLcs.AI2025-11AAAI被引 3

构建5400亿词的印地语族合成数据集,提升低资源语言模型性能

BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages

  • 用5种技术生成跨10种印地语族语言的合成数据
  • 540B tokens数据集在模型训练中显著提升低资源语言表现
  • 支持多语言、多脚本的质量评估框架,适合语言研究者和NLP工程师

在大型语言模型预训练背景下,合成数据成为大规模生成高质量预训练数据的替代方案,尤其对低资源语言尤为关键。本文系统研究了印地语族语言的多语言合成预训练数据生成与评估,构建了包含5400亿词的BhashaKritika大规模合成数据集,涵盖10种语言,采用5种不同技术。研究探索了基于文档、人物设定和主题的生成方式影响,分析了提示词与文档底层语言选择对数据质量的影响,并对比了英文内容翻译与原生印地语族生成的效果。为支持可扩展且语言敏感的评估,提出模块化质量评估流程,集成脚本与语言检测、元数据一致性检查、n-gram重复分析及基于KenLM模型的困惑度过滤。实验结果揭示生成策略的关键权衡,明确了构建高效多语言语料库的最佳实践。

原文摘要 · Abstract (English)

In the context of pretraining of Large Language Models (LLMs), synthetic data has emerged as an alternative for generating high-quality pretraining data at scale. This is particularly beneficial in low-resource language settings where the benefits of recent LLMs have been unevenly distributed across languages. In this work, we present a systematic study on the generation and evaluation of synthetic multilingual pretraining data for Indic languages, where we construct a large-scale synthetic dataset BhashaKritika, comprising 540B tokens using 5 different techniques for 10 languages. We explore the impact of grounding generation in documents, personas, and topics. We analyze how language choice, both in the prompt instructions and document grounding, affects data quality, and we compare translations of English content with native generation in Indic languages. To support scalable and language-sensitive evaluation, we introduce a modular quality evaluation pipeline that integrates script and language detection, metadata consistency checks, n-gram repetition analysis, and perplexity-based filtering using KenLM models. Our framework enables robust quality control across diverse scripts and linguistic contexts. Empirical results through model runs reveal key trade-offs in generation strategies and highlight best practices for constructing effective multilingual corpora.

合成数据印地语族LLM预训练多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。