arXiv:2509.23673cs.CVcs.AI2025-09EMNLP被引 4

提出RCI评分,量化多模态模型是依赖全局理解还是局部线索。

RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks

  • 通过对比模型在图像块与全图上的表现,区分任务对全局或局部视觉信息的依赖。
  • 13个主流数据集测试显示多数依赖局部线索,存在显著空间偏差。
  • 为构建更鲁棒的多模态系统提供可操作的诊断工具,适合研究者与开发者使用。

多模态大语言模型在视觉-语言基准上取得显著进展,但现有基准是否真正评估全局推理能力仍不明确,部分结果可能仅依赖局部视觉线索。现有评估方法无法明确区分这一差异,阻碍了数据集优化和面向实际应用的模型开发。本文提出区域理解指数(RCI),首个基于模型的评分体系,直接量化数据集对全局与局部视觉信息的依赖程度。RCI系统性比较参考模型在图像块与完整图像上的表现,揭示任务是否需要整体图像理解,或仅靠局部视觉线索即可解决。将RCI应用于13个广泛使用的多模态基准发现,多数任务偏好局部推理,且存在显著空间偏差,暗示其在真实场景中的潜在风险。RCI为研究人员与实践者提供了可操作的工具,用于诊断并缓解此类偏差,助力构建更具鲁棒性的企业级多模态系统。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved impressive results on vision-language benchmarks, yet it remains unclear whether these benchmarks assess genuine global reasoning or allow success via localized visual cues. Existing evaluation methods do not explicitly measure this distinction, hindering effective dataset curation and real-world focused model development. We introduce Region Comprehension Index (RCI), the first model-based score to directly quantify a dataset's reliance on global versus local visual information. RCI systematically compares reference-model performance on image patches versus full images, revealing if tasks require holistic image understanding or can be solved with partial or localized visual cues. When applying RCI to 13 widely used multimodal benchmarks, we observed that most of them favor localized reasoning and exhibit significant spatial biases, indicating potential risks in real-world applications. RCI equips researchers & practitioners with an actionable tool for diagnosing & mitigating these biases, enabling the construction of datasets and benchmarks to foster the development of robust, enterprise-ready multimodal systems.

多模态评估指标模型偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。