为印度语言打造高质量指令微调数据集,解决文化语境缺失问题
Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
- 人机协同流水线融合翻译与合成生成,保障数据多样性
- 构建22.5K条指令数据(Pragyaan-IT)与100K条对齐数据(Pragyaan-Align)
- 覆盖10种印度语言,强调文化细节与多轮对话的真实还原
大语言模型(LLM)的效果高度依赖高质量的后训练数据,尤其是指令微调和偏好对齐样本。现有开源数据集普遍存在多语言覆盖不足、文化语境缺失及任务多样性不足的问题,尤其在印度语言中更为显著。我们提出一种人机协同的流水线方法,结合翻译与合成扩展,生成可靠且多样化的印地语系后训练数据。基于此流程,我们构建了两个数据集:Pragyaan-IT(22.5K条)和Pragyaan-Align(100K条),覆盖10种印度语言,涵盖13个主类别与56个子类别,整合了57个多样化数据源。数据协议包含常被忽视的维度,强调任务多样性、多轮对话、指令忠实性、安全对齐及文化细微差别的保留,为更包容、高效的多语言大模型提供基础。
原文摘要 · Abstract (English)
The effectiveness of Large Language Models (LLMs) depends heavily on the availability of high-quality post-training data, particularly instruction-tuning and preference-based examples. Existing open-source datasets, however, often lack multilingual coverage, cultural grounding, and suffer from task diversity gaps that are especially pronounced for Indian languages. We introduce a human-in-the-loop pipeline that combines translations with synthetic expansion to produce reliable and diverse Indic post-training data. Using this pipeline, we curate two datasets: Pragyaan-IT (22.5K) and Pragyaan-Align (100K) across 10 Indian languages covering 13 broad and 56 sub-categories, leveraging 57 diverse datasets. Our dataset protocol incorporates several often-overlooked dimensions and emphasize task diversity, multi-turn dialogue, instruction fidelity, safety alignment, and preservation of cultural nuance, providing a foundation for more inclusive and effective multilingual LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。