arXiv:2509.23499cs.CVcs.CL2025-09中稿 · ICLR被引 3

揭示多模态数据中视觉与文本依赖的复杂关系,指出当前基准存在偏倚放大问题。

Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional

  • 用23个VQA基准测试多模态大模型,量化各模态独立与交互贡献。
  • 发现多数基准反而强化了图像依赖,模型常独立使用模态而非融合交互。
  • 为多模态评估提供可量化的基准设计准则,适合研究者与评测人员参考。

理解单模态依赖(单个模态对任务的贡献)与跨模态依赖(模态间及与目标任务的关系)的相互作用,是推进多模态学习的关键。然而,当前基准评估中这些依赖的性质与交互方式仍不清晰。本文通过大规模实证研究,基于覆盖通用与专家知识推理、光学字符识别、文档理解等领域的多模态大语言模型(MLLMs),在23个视觉问答(VQA)基准上量化了这些依赖。结果表明,对视觉、问题(文本)及其交互的依赖在不同及同一基准间差异显著。我们发现,许多旨在缓解纯文本偏见的基准反而意外增强了纯图像依赖。这种现象在不同模型规模与类型中均存在,模型通常通过独立使用各模态获得高表现,而对模态间交互依赖有限。本研究提供了多模态数据集的定量刻画,支持更严谨的多模态基准设计与评估。

原文摘要 · Abstract (English)

Understanding the interplay between intra-modality dependencies (the contribution of an individual modality to a target task) and inter-modality dependencies (the relationships between modalities and the target task) is fundamental to advancing multi-modal learning. However, the nature of and interaction between these dependencies within current benchmark evaluations remains poorly characterized. In this work, we present a large-scale empirical study to quantify these dependencies across 23 visual question-answering benchmarks using multi-modal large language models (MLLMs) covering domains such as general and expert knowledge reasoning, optical character recognition, and document understanding. Our findings show that the reliance on vision, question (text), and their interaction varies significantly, both across and within benchmarks. We discover that numerous benchmarks intended to mitigate text-only biases have inadvertently amplified image-only dependencies. This characterization persists across model sizes and types, with models often obtaining high performance by using each modality independently and showing limited dependence on their interaction. We provide a quantitative characterization of multi-modal datasets, enabling a principled approach to multi-modal benchmark design and evaluation.

多模态基准评测视觉问答模型依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。