评测大模型跨语言视觉问答能力,发现中文等语言存在输出偏差
LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
- 构建11种语言的多模态评测基准,覆盖5类社会属性
- 闭源模型整体表现最优,但各语言间公平性仍不均衡
- 开源模型如Qwen2.5展现良好多语言泛化能力,适合跨语言研究
大型多模态模型(LMMs)通常在海量图文数据上训练,但语言覆盖有限,导致跨语言输出存在偏见与不公平。现有研究多关注多模态评估,却较少关注多语言能力。本文提出LinguaMark基准,用于评估前沿LMMs在多语言视觉问答(VQA)任务中的表现。该数据集包含6,875个跨语言图像-文本对,涵盖11种语言和5类社会属性。采用偏见(Bias)、答案相关性(Answer Relevancy)和忠实度(Faithfulness)三项核心指标进行评估。结果表明,闭源模型(GPT-4o、Gemini2.5)整体性能最高;闭源与开源模型(Gemma3、Qwen2.5)在社会属性任务中表现相当,其中Qwen2.5展现出优异的多语言泛化能力。研究代码与数据集已公开,以促进可复现性与后续研究。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) are typically trained on vast corpora of image-text data but are often limited in linguistic coverage, leading to biased and unfair outputs across languages. While prior work has explored multimodal evaluation, less emphasis has been placed on assessing multilingual capabilities. In this work, we introduce LinguaMark, a benchmark designed to evaluate state-of-the-art LMMs on a multilingual Visual Question Answering (VQA) task. Our dataset comprises 6,875 image-text pairs spanning 11 languages and five social attributes. We evaluate models using three key metrics: Bias, Answer Relevancy, and Faithfulness. Our findings reveal that closed-source models generally achieve the highest overall performance. Both closed-source (GPT-4o and Gemini2.5) and open-source models (Gemma3, Qwen2.5) perform competitively across social attributes, and Qwen2.5 demonstrates strong generalization across multiple languages. We release our benchmark and evaluation code to encourage reproducibility and further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。