arXiv:2604.19405cs.CL2026-04中稿 · ACL

首个多语言视觉-语言评估基准,揭示大模型评判能力跨语言差异

Lost in Translation: Do LVLM Judges Generalize Across Languages?

论文配图:Lost in Translation: Do LVLM Judges Generalize Across Languages?
图 1 · 摘自论文原文
  • 构建6万+跨语言配对偏好数据集,覆盖25种语言类型
  • 发现主流模型在不同语言中评分表现差异巨大,性能不可靠
  • 适合关注多语言模型评估与对齐的研究者参考

自动评价器如奖励模型在大型视觉-语言模型(LVLM)对齐与评估中扮演核心角色。然而,这些评价器几乎仅在以英语为中心的基准上被评估,其跨语言泛化能力仍不明确。为此,我们提出MM-JudgeBench,首个大规模多语言、多模态评价基准,包含超过6万对跨语言偏好实例,覆盖25种语言类型。该基准包含两个互补子集:扩展自VL-RewardBench的通用视觉-语言偏好评估子集,以及源自OpenCQA的图表中心视觉-文本推理子集,支持对奖励模型(即LVLM评判者)在多种场景下的系统分析。我们还发布了独立于评估数据的多语言训练集(来自MM-RewardBench),用于领域适应。通过对22个LVLM(15个开源,7个专有)的评估,我们发现其在跨语言情境下表现存在显著差异。分析表明,模型规模与架构无法有效预测多语言鲁棒性,即使最先进的LVLM评判者在不同语言间行为也不一致。这些发现揭示了当前奖励建模的根本局限,并强调建立多语言、多模态基准对开发可靠自动化评价器的必要性。

原文摘要 · Abstract (English)

Automatic evaluators such as reward models play a central role in the alignment and evaluation of large vision-language models (LVLMs). Despite their growing importance, these evaluators are almost exclusively assessed on English-centric benchmarks, leaving open the question of how well these evaluators generalize across languages. To answer this question, we introduce MM-JudgeBench, the first large-scale benchmark for multilingual and multimodal judge model evaluation, which includes over 60K pairwise preference instances spanning 25 typologically diverse languages. MM-JudgeBench integrates two complementary subsets: a general vision-language preference evaluation subset extending VL-RewardBench, and a chart-centric visual-text reasoning subset derived from OpenCQA, enabling systematic analysis of reward models (i.e., LVLM judges) across diverse settings. We additionally release a multilingual training set derived from MM-RewardBench, disjoint from our evaluation data, to support domain adaptation. By evaluating 22 LVLMs (15 open-source, 7 proprietary), we uncover substantial cross-lingual performance variance in our proposed benchmark. Our analysis further shows that model size and architecture are poor predictors of multilingual robustness, and that even state-of-the-art LVLM judges exhibit inconsistent behavior across languages. Together, these findings expose fundamental limitations of current reward modeling and underscore the necessity of multilingual, multimodal benchmarks for developing reliable automated evaluators.

多语言模型评估视觉语言奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。