arXiv:2508.05502cs.CVcs.AI2025-08被引 1

构建多语言图文数据集,提升低资源语言的跨文化理解能力

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

  • 用本土图文对与生成翻译对分离监督信号
  • 在8种低资源语言中减少文化幻觉现象
  • 适合关注多语言视觉认知的科研人员

多模态大模型在高资源语言中表现良好,但在低资源语言中常产生流利却缺乏文化内涵的描述。我们认为问题不仅在于语言能力:文化特定视觉知识依赖于母语的视觉-文本对齐,而以翻译为中心的流程难以提供此类对齐。为此,我们提出MELLA,一个覆盖八种低资源语言的多模态数据集,旨在同时支持语言流畅性与文化扎根性。MELLA采用双源策略:利用本土网络图像与替代文本对提供文化相关监督,结合生成并翻译的图像描述提供语言丰富监督,明确分离通常混杂在多语言多模态数据中的两类学习信号。在多个MLLM主干模型上进行受控诊断微调后,我们发现MELLA通过帮助模型识别和表达被基于翻译的适配忽略的文化特异性实体,有效缓解了文化幻觉。研究强调,数据对齐而非仅模型修改,是实现低资源语言中文化扎根多模态理解的关键路径。数据集已开放获取:https://opendatalab.com/applyMultilingualCorpus。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings. We argue that this failure is not merely a linguistic limitation: culture-specific visual knowledge depends on native visual-textual alignments that translation-centric pipelines rarely provide. We present MELLA, a multimodal dataset across eight low-resource languages, designed to support linguistic fluency and cultural groundedness. MELLA uses a dual-source strategy that combines native web image-alt-text pairs for culture-grounded supervision with generated-and-translated image descriptions for linguistically rich supervision, explicitly separating two learning signals often conflated in multilingual multimodal data. Through controlled diagnostic fine-tuning on multiple MLLM backbones, we show that MELLA mitigates cultural hallucination by helping models recognize and articulate culturally specific entities overlooked by translation-based adaptation. Our findings highlight data alignment, rather than model modification alone, as a path toward culturally grounded multimodal understanding in low-resource languages. The dataset is available at https://opendatalab.com/applyMultilingualCorpus.

多模态模型低资源语言文化对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。