arXiv:2502.05568cs.CLcs.AI2025-02中稿 · Information Fusion综述被引 19

梳理117项研究,帮低资源语言用上大模型多模态技术。

Large Multimodal Models for Low-Resource Languages: A Survey

  • 按数据和方法分类,系统整理适配低资源语言的多模态模型策略。
  • 发现视觉信息能显著提升低资源语言模型表现,但幻觉和效率仍存挑战。
  • 适合关注多模态模型本地化、跨语言研究的研究者参考。

本综述系统分析了将大模型多模态(LMMs)适配至低资源(LR)语言的技术,涵盖视觉增强、数据生成、跨模态迁移与融合策略。通过对96种低资源语言的117项研究进行综合分析,识别出研究者应对数据与计算资源匮乏的关键模式。研究按资源导向与方法导向分类,并细分为子类别。比较方法导向研究在性能与效率上的表现,讨论代表性工作的优劣。结果表明,视觉信息常作为提升低资源场景下模型表现的关键桥梁,但幻觉抑制与计算效率仍面临显著挑战。本文为研究人员提供了当前进展与未解决问题的清晰图景,助力推动多模态模型对低资源语言使用者的可及性。附开源仓库:https://github.com/marianlupascu/LMM4LRL-Survey。

原文摘要 · Abstract (English)

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) for low-resource (LR) languages, examining approaches ranging from visual enhancement and data creation to cross-modal transfer and fusion strategies. Through a comprehensive analysis of 117 studies across 96 LR languages, we identify key patterns in how researchers tackle the challenges of limited data and computational resources. We categorize works into resource-oriented and method-oriented contributions, further dividing contributions into relevant sub-categories. We compare method-oriented contributions in terms of performance and efficiency, discussing benefits and limitations of representative studies. We find that visual information often serves as a crucial bridge for improving model performance in LR settings, though significant challenges remain in areas such as hallucination mitigation and computational efficiency. In summary, we provide researchers with a clear understanding of current approaches and remaining challenges in making LMMs more accessible to speakers of LR (understudied) languages. We complement our survey with an open-source repository available at: https://github.com/marianlupascu/LMM4LRL-Survey.

多模态低资源语言大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。