arXiv:2601.10388cs.CL2026-01

构建首个印地语方言多任务评测基准,助力低资源方言研究

INDIC DIALECT: A Multi Task Benchmark to Evaluate and Translate in Indian Language Dialects

  • 构建13k句对的跨方言平行语料库,覆盖11种方言与2种语言
  • 细调模型使方言分类F1从19.6%提升至89.8%,翻译性能显著改善
  • 提出混合生成策略,规则+AI在方言生成中表现最优

当前自然语言处理进展主要集中于标准语言,导致多数低资源方言未被充分关注,尤其在印度尤为突出。尽管印地语是全球第三大使用语言(超6亿使用者),其众多方言仍缺乏代表性;奥里亚语也有约4500万使用者。现有数据集虽包含标准印地语和奥里亚语,但地区方言几乎无网络内容。我们提出INDIC-DIALECT,一个由人工标注的平行语料库,包含13,000句对,覆盖11种方言与2种语言:印地语和奥里亚语。基于此,我们构建了一个包含三任务的多任务基准:方言分类、多选题问答与机器翻译。实验表明,如GPT-4o和Gemini 2.5等大模型在分类任务上表现不佳;而基于印度语言预训练的微调变换器模型显著提升性能,例如将方言分类的F1从19.6%提升至89.8%。在方言到标准语翻译中,混合人工智能模型达到最高BLEU得分61.32,远超基线23.36。值得注意的是,在标准语到方言翻译中,采用“规则引导+AI生成”的方法取得最佳BLEU得分48.44,优于基线27.59。INDIC-DIALECT为印地语族方言感知的NLP研究提供新基准,计划开源以推动低资源方言研究。

原文摘要 · Abstract (English)

Recent NLP advances focus primarily on standardized languages, leaving most low-resource dialects under-served especially in Indian scenarios. In India, the issue is particularly important: despite Hindi being the third most spoken language globally (over 600 million speakers), its numerous dialects remain underrepresented. The situation is similar for Odia, which has around 45 million speakers. While some datasets exist which contain standard Hindi and Odia languages, their regional dialects have almost no web presence. We introduce INDIC-DIALECT, a human-curated parallel corpus of 13k sentence pairs spanning 11 dialects and 2 languages: Hindi and Odia. Using this corpus, we construct a multi-task benchmark with three tasks: dialect classification, multiple-choice question (MCQ) answering, and machine translation (MT). Our experiments show that LLMs like GPT-4o and Gemini 2.5 perform poorly on the classification task. While fine-tuned transformer based models pretrained on Indian languages substantially improve performance e.g., improving F1 from 19.6\% to 89.8\% on dialect classification. For dialect to language translation, we find that hybrid AI model achieves highest BLEU score of 61.32 compared to the baseline score of 23.36. Interestingly, due to complexity in generating dialect sentences, we observe that for language to dialect translation the ``rule-based followed by AI" approach achieves best BLEU score of 48.44 compared to the baseline score of 27.59. INDIC-DIALECT thus is a new benchmark for dialect-aware Indic NLP, and we plan to release it as open source to support further work on low-resource Indian dialects.

方言识别机器翻译多任务学习低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。