arXiv:2603.13427cs.CVcs.AI2026-03

构建多模态交互评估基准,揭示大模型在跨模态融合中的能力瓶颈。

MIBench: Evaluating LMMs on Multimodal Interaction

  • 设计三类认知层级的多模态任务,测试模型从识别到推理的交互能力
  • 10,000+数据对验证主流模型仍受限于模态干扰与协同生成能力弱
  • 适合研究多模态模型泛化与鲁棒性的开发者参考

在不同多模态场景中,需根据任务需求以特定方式整合与利用视觉和文本信息,这种整合方式称为“多模态交互”。模型处理多模态交互的能力直接反映其多模态理解水平。本文提出MIBench,一个全面评估大尺寸多模态模型(LMMs)多模态交互能力的基准。每个实例被建模为(con_v, con_t, task)三元组,包含视觉与文本上下文,要求模型采用正确的多模态交互形式完成任务。MIBench从三个关键维度评估:基于视觉或文本线索的信息获取能力,以及二者联合产生的新信息生成能力。每项能力在三个认知层级(识别、理解、推理)上分层评估。该基准包含超过10,000个视觉-文本上下文对,覆盖32种任务。对先进LMMs的评估显示:(1) 尽管模型参数与训练数据规模持续增长,其多模态交互能力仍受限制;(2) 在处理视觉信息时易被文本模态干扰;(3) 多模态协同生成能力普遍较弱;(4) 原生训练的多模态模型在基础交互能力上存在明显不足。这些发现可为未来提升模型多模态能力提供参考。

原文摘要 · Abstract (English)

In different multimodal scenarios, it needs to integrate and utilize information across modalities in a specific way based on the demands of the task. Different integration ways between modalities are referred to as "multimodal interaction". How well a model handles various multimodal interactions largely characterizes its multimodal ability. In this paper, we introduce MIBench, a comprehensive benchmark designed to evaluate the multimodal interaction capabilities of Large Multimodal Models (LMMs), which formulates each instance as a (con_v , con_t, task) triplet with contexts from vision and text, necessitating that LMMs employ correct forms of multimodal interaction to effectively complete the task. MIBench assesses models from three key aspects: the ability to source information from vision-centric or text-centric cues, and the ability to generate new information from their joint synergy. Each interaction capability is evaluated hierarchically across three cognitive levels: Recognition, Understanding, and Reasoning. MIBench comprises over 10,000 vision-text context pairs spanning 32 distinct tasks. Evaluation of state-of-the-art LMMs show that: (1) LMMs' ability on multimodal interaction remains constrained, despite the scaling of model parameters and training data; (2) they are easily distracted by textual modalities when processing vision information; (3) they mostly possess a basic capacity for multimodal synergy; and (4) natively trained multimodal models show noticeable deficits in fundamental interaction ability. We expect that these observations can serve as a reference for developing LMMs with more enhanced multimodal ability in the future.

多模态评估交互能力大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。