arXiv:2604.18423cs.CL2026-04ACL综述被引 1

首份整合印度语NLP资源的综述,覆盖200+数据集和100+模型。

BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources

论文配图:BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources
图 1 · 摘自论文原文
  • 按语言现象、领域和模态分类整理印度语NLP资源
  • 涵盖200+数据集、50+评测基准、100+模型与工具
  • 适合关注低资源语言与文化适配NLP的研究者

印度拥有22种官方语言及数百种边缘化方言,推动了自然语言处理(NLP)数据集、基准和预训练模型的快速增长。然而,尚无专门针对印度语言的系统性综述。现有研究或聚焦少数高资源语言,或将其纳入更广泛的多语言框架中,难以覆盖低资源和文化多样性语言。为此,本文首次提供一份统一的印度NLP资源综述,涵盖200多个数据集、50多个评测基准以及100多个模型、工具与系统,覆盖文本、语音、多模态和文化相关任务。按语言现象、领域和模态组织资源,分析标注、评估与模型设计趋势,并识别出数据稀缺、语言覆盖不均、文字多样性及文化与领域泛化能力不足等持续挑战。本综述为实现公平、文化适配且可扩展的印度语NLP研究提供了整合基础。

原文摘要 · Abstract (English)

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates resources developed specifically for Indian languages. Existing reviews either focus on a few high-resource languages or subsume Indian languages within broader multilingual settings, limiting coverage of low-resource and culturally diverse varieties. To address this gap, we present the first unified survey of Indian NLP resources, covering 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks. We organize resources by linguistic phenomena, domains, and modalities; analyze trends in annotation, evaluation, and model design; and identify persistent challenges such as data sparsity, uneven language coverage, script diversity, and limited cultural and domain generalization. This survey offers a consolidated foundation for equitable, culturally grounded, and scalable NLP research in the Indian linguistic ecosystem.

印度语言数据集综述多语言NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。