测试开源多模态模型在反常识画面中的判断力,发现它们更信文字而非眼睛。
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
- 用400个反常识图像测试模型,如兔子追老虎
- 开源模型准确率仅随机水平,远低于人类和闭源模型
- 通过微调和结构化提示可让模型信任视觉证据
多模态大语言模型在主流视觉理解任务中表现优异,但在违背日常常识的动作场景中能力尚未充分验证。为此,我们构建了CAIT基准,包含400个高保真合成场景,聚焦反常识视觉动作,如“兔子追老虎”,视觉证据明确与常识冲突。我们评估了人类、领先闭源模型(如Claude和Gemini)以及14个代表性开源模型。人类表现接近完美(约0.95准确率),闭源模型表现出稳健理解(最高达0.88准确率),而标准开源指令微调模型表现仅处于随机水平。进一步分析表明,这种失败源于强烈的语言先验:模型不信任异常视觉输入,而是自动以统计上常见的文本描述覆盖视觉信号。尽管引入思维链机制可提升准确率,但显著降低响应速度,并引发新问题:模型过度思考,因违反物理法则而拒绝接受真实视觉内容。最后,我们证明针对性微调和结构化提示能有效缓解对语言先验的依赖,使开源模型准确基于实际视觉证据进行推理。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in mainstream visual understanding tasks, but their ability to process action scenes that contradict everyday common sense remains undertested. To address this gap, we introduce CAIT, a benchmark comprising 400 high-fidelity synthetic scenes focused on counter-intuitive visual actions, such as ``a rabbit is chasing a tiger'', where visual evidence explicitly contradicts common-sense expectations. We evaluate human, leading proprietary models (e.g., Claude and Gemini), and 14 representative open-source MLLMs. Humans achieve near-perfect performance (around 0.95 accuracy) and proprietary models demonstrate robust understanding (achieving up to 0.88 accuracy), standard open-source instruction-tuned models perform at the chance level. Further analysis demonstrates that this failure is driven by a strong language prior: rather than trusting the visual input, they automatically override the anomalous visual signals with statistically common text descriptions. Although introducing Chain-of-Thought reasoning mechanisms can improve accuracy, it significantly slows down the response and generates a new failure mode: models overthink the scenario and refuse to accept the actual visual content simply because it violates real-world physical laws. Finally, we demonstrate that targeted fine-tuning and structured prompting can effectively mitigate this reliance on language priors, enabling open-source models to accurately ground their reasoning in actual visual evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。