梳理南亚低资源语言NLP现状,揭示数据与评测短板
Bhaasha, Bhasa, Zaban: A Survey for Low-Resourced Languages in South Asia -- Current Stage and Challenges
- 系统回顾2020年后南亚语言Transformer模型研究
- 发现健康等关键领域数据缺失、缺乏统一评测基准
- 适合关注多元语言公平与本地化NLP的研究者
大型语言模型的快速发展已显著推动英语自然语言处理任务,但南亚地区众多低资源语言仍被忽视。尽管南亚有超过650种语言,许多语言要么计算资源极少,要么未被现有模型覆盖。本综述通过检索2020年以来的相关研究,聚焦基于Transformer的模型(如BERT、T5、GPT),全面分析南亚语言在数据、模型与任务三个核心维度的进展与挑战,涵盖可用数据源、微调策略及应用场景。研究发现存在严重问题:关键领域(如医疗)数据缺失、语言混用现象普遍、缺乏标准化评估基准。本文旨在提升学术界对南亚语言公平性的关注,推动针对性数据建设、建立契合文化语境的统一评测体系,并促进南亚语言在主流NLP中的均衡代表。完整资源列表见:https://github.com/trust-nlp/LM4SouthAsia-Survey。
原文摘要 · Abstract (English)
Rapid developments of large language models have revolutionized many NLP tasks for English data. Unfortunately, the models and their evaluations for low-resource languages are being overlooked, especially for languages in South Asia. Although there are more than 650 languages in South Asia, many of them either have very limited computational resources or are missing from existing language models. Thus, a concrete question to be answered is: Can we assess the current stage and challenges to inform our NLP community and facilitate model developments for South Asian languages? In this survey, we have comprehensively examined current efforts and challenges of NLP models for South Asian languages by retrieving studies since 2020, with a focus on transformer-based models, such as BERT, T5, & GPT. We present advances and gaps across 3 essential aspects: data, models, & tasks, such as available data sources, fine-tuning strategies, & domain applications. Our findings highlight substantial issues, including missing data in critical domains (e.g., health), code-mixing, and lack of standardized evaluation benchmarks. Our survey aims to raise awareness within the NLP community for more targeted data curation, unify benchmarks tailored to cultural and linguistic nuances of South Asia, and encourage an equitable representation of South Asian languages. The complete list of resources is available at: https://github.com/trust-nlp/LM4SouthAsia-Survey.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。