综述多语言视觉-语言模型,揭示中立与文化适应的矛盾
Multilingual Vision-Language Models, A Survey
- 分析33个模型23个基准,对比编码器与生成式架构
- 三分之二评测依赖翻译对齐,忽视文化语境差异
- 指出训练目标与评估标准不一致,影响跨语言性能
本综述系统考察了处理多语言文本与图像的视觉-语言模型。涵盖33个模型和23个基准,涉及仅编码器与生成式架构,识别出语言中立性(跨语言表示一致性)与文化敏感性(适配文化背景)之间的核心矛盾。当前训练方法多通过对比学习强化中立性,而文化适应依赖多样数据。三分之二的评估基准采用基于翻译的方法,强调语义一致性,但近期研究开始引入文化相关内容。发现跨语言能力存在差异,且训练目标与评估目标之间存在明显脱节。
原文摘要 · Abstract (English)
This survey examines multilingual vision-language models that process text and images across languages. We review 33 models and 23 benchmarks, spanning encoder-only and generative architectures, and identify a key tension between language neutrality (consistent cross-lingual representations) and cultural awareness (adaptation to cultural contexts). Current training methods favor neutrality through contrastive learning, while cultural awareness depends on diverse data. Two-thirds of evaluation benchmarks use translation-based approaches prioritizing semantic consistency, though recent work incorporates culturally grounded content. We find discrepancies in cross-lingual capabilities and gaps between training objectives and evaluation goals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。