arXiv:2505.04531cs.CLcs.AI2025-05综述被引 17

系统梳理低资源语言生成模型的数据稀缺问题与解决方案

Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review

  • 综述54项研究,分类整理数据增强、回译等技术策略
  • 发现主流依赖Transformer模型,仅覆盖少数低资源语言
  • 呼吁统一评估标准,推动更具包容性的语言AI发展

生成式语言建模因ChatGPT、Google Gemini等服务兴起而备受关注,但其主要服务于英语等高资源语言,加剧了自然语言处理领域的语言不平等。本文首次针对低资源语言(LRL)生成式建模中的数据稀缺问题进行系统性综述。基于54项研究,我们识别、分类并评估了单语数据增强、回译、多语言训练和提示工程等技术方法在生成任务中的应用。同时分析了模型架构选择、语系覆盖范围及评估方法的趋势。研究发现,现有工作高度依赖Transformer模型,主要集中于少数低资源语言,且评估方式缺乏一致性。最后提出扩展方法覆盖更广泛低资源语言的建议,并指出构建公平生成式语言系统面临的开放挑战。本综述旨在支持研究人员与开发者打造面向弱势语言的包容性AI工具,是实现语言多样性保护的关键一步。

原文摘要 · Abstract (English)

Generative language modelling has surged in popularity with the emergence of services such as ChatGPT and Google Gemini. While these models have demonstrated transformative potential in productivity and communication, they overwhelmingly cater to high-resource languages like English. This has amplified concerns over linguistic inequality in natural language processing (NLP). This paper presents the first systematic review focused specifically on strategies to address data scarcity in generative language modelling for low-resource languages (LRL). Drawing from 54 studies, we identify, categorise and evaluate technical approaches, including monolingual data augmentation, back-translation, multilingual training, and prompt engineering, across generative tasks. We also analyse trends in architecture choices, language family representation, and evaluation methods. Our findings highlight a strong reliance on transformer-based models, a concentration on a small subset of LRLs, and a lack of consistent evaluation across studies. We conclude with recommendations for extending these methods to a wider range of LRLs and outline open challenges in building equitable generative language systems. Ultimately, this review aims to support researchers and developers in building inclusive AI tools for underrepresented languages, a necessary step toward empowering LRL speakers and the preservation of linguistic diversity in a world increasingly shaped by large-scale language technologies.

低资源语言生成模型系统综述语言公平

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。