arXiv:2509.21294cs.CL2025-09ACL被引 2

为13种印度语言构建高质量指令数据,提升多语言AI文化适配性

UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languages

  • 用本土维基内容生成指令数据,避免机械翻译偏差
  • 覆盖950万条数据,模型在13种语言上性能显著提升
  • 适合多语言低资源场景下的AI训练与评估

开发文化契合的多语言AI系统仍具挑战,尤其对低资源语言。现有合成数据多依赖英语翻译,缺乏文化语境。本文提出基于大模型(≥2350亿参数)与本地维基内容的自下而上合成方法,构建高质大型指令跟随数据集Updesh,包含13种印度语言及英语共950万条数据,涵盖多样推理与生成任务。通过自动化指标与1万次人工评估验证数据质量。下游微调实验表明,基于Updesh训练的模型在13个多样化多语言数据集上的自然语言理解(NLU)与生成(NLG)任务中均取得显著提升。消融与文化评估进一步证明,上下文敏感、文化扎根的数据生成对多语言AI发展至关重要。

原文摘要 · Abstract (English)

Developing culturally grounded multilingual AI systems remains challenging, particularly for low-resource languages. While synthetic data offers promise, its effectiveness in multilingual and multicultural contexts is underexplored. We investigate bottom-up synthetic data generation using large open-source LLMs (>= 235B parameters) grounded in language-specific Wikipedia content, complementing dominant top-down translation-based approaches from English. We introduce Updesh, a high-quality large-scale synthetic instruction-following dataset comprising 9.5M data points across 13 Indian languages and English, encompassing diverse reasoning and generative tasks. Comprehensive evaluation using automated metrics and 10K human assessments confirms high data quality. Downstream evaluations performed by fine-tuning models on various datasets and assessing performance across 13 diverse multilingual datasets and model comparative evaluations, demonstrate that models trained on Updesh consistently obtain significant improvements on NLU, NLG evaluations. Finally, through ablation studies and cultural evaluations, we show that context-aware, culturally grounded data generation is essential for effective multilingual AI development.

多语言指令数据文化适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。