arXiv:2608.22232cs.AIcs.CL2026-08

揭示多模态大模型在情境错觉下的脆弱性并提出改进方法

Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models

论文配图:Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models
图 1 · 摘自论文原文
  • 构建了'何处-何物-如何'分类体系,系统刻画情境错觉
  • 27个模型测试显示普遍易受误导,存在6类典型失败模式
  • 提出提示工程与微调方案,性能提升最高达20%

现实情境的表象常与其物理本质不符,挑战多模态大语言模型(MLLMs)在实际应用中的可靠性。本文将此现象称为情境错觉,并研究:(1)MLLMs在该类错觉下的表现,(2)如何缓解其局限。我们首先提出一个全面的where-what-how分类体系,用于描述情境错觉发生的地点、目标及其生成机制。基于此,构建了MSIBench基准,用于评估MLLMs在情境错觉下的辨别、理解与推理能力。对27种模型配置的评估显示,当前MLLMs高度易受此类错觉影响,表现出与视觉观察、语义定位及推理相关的6类典型失败模式。为缓解缺陷,我们提出两种方法:针对闭源模型采用基于视觉证据的系统性提示策略;针对开源模型则采用监督微调。这两种简单但有效的手段使模型性能提升最多达20%,为复杂真实环境中的可靠多模态感知与推理提供了可行路径。

原文摘要 · Abstract (English)

Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large language models (MLLMs) in practical applications. In this paper, we term this phenomenon situational illusions and investigate: (1) how MLLMs perform under such illusions, and (2) how to mitigate the limitations. We first develop a comprehensive where-what-how taxonomy that characterizes where situational illusions occur, what targets they take, and how they arise. Building on this taxonomy, we introduce MSIBench, a benchmark designed to assess the discrimination, understanding, and reasoning capabilities of MLLMs under situational illusions. Evaluations of 27 model configurations reveal that current MLLMs are highly vulnerable to these illusions and exhibit 6 typical failure modes related to visual observation, grounding, and reasoning. To mitigate the limitations, we build on the core idea of systematically inspecting and reasoning over visual evidence for contextual understanding, developing prompting for closed-source models and supervised fine-tuning for open-source models, respectively. These two simple yet effective methods improve model performances by 20% at most, suggesting a practical path toward more reliable multimodal perception and reasoning in complex real-world environments.

多模态大模型情境错觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。