arXiv:2411.14137cs.CVcs.CL2024-11ICCV被引 5

用视觉上下文解构模糊表达,评测模型真实意图理解能力

VAGUE: Visual Contexts Clarify Ambiguous Expressions

  • 构建1.6万条图文配对的模糊语义测试集,需结合图像才能判断真意
  • 现有模型准确率远低于人类,即使增加视觉线索提升有限
  • 模型常误认表面关联而非真正意图,缺乏深层多模态推理

人类交流常依赖视觉线索化解歧义,但当前AI在多模态推理上仍显不足。本文提出VAGUE基准,包含1.6K条需结合图像才能理解的模糊文本表达,每条配有图像和多项选择解释,正确答案仅在视觉上下文中显现。数据覆盖人工构造的复杂场景(Visual Commonsense Reasoning)与自然真实的个人视角场景(Ego4D),确保多样性。实验表明,现有多模态模型难以准确推断说话者真实意图;虽引入更多视觉线索后性能有所提升,整体准确率仍显著低于人类水平,暴露出关键的多模态推理差距。失败案例分析显示,当前模型往往捕捉视觉场景中的表面相关性而非真实意图,说明其‘看到’图像却未有效‘理解’。

原文摘要 · Abstract (English)

Human communication often relies on visual cues to resolve ambiguity. While humans can intuitively integrate these cues, AI systems often find it challenging to engage in sophisticated multimodal reasoning. We introduce VAGUE, a benchmark evaluating multimodal AI systems' ability to integrate visual context for intent disambiguation. VAGUE consists of 1.6K ambiguous textual expressions, each paired with an image and multiple-choice interpretations, where the correct answer is only apparent with visual context. The dataset spans both staged, complex (Visual Commonsense Reasoning) and natural, personal (Ego4D) scenes, ensuring diversity. Our experiments reveal that existing multimodal AI models struggle to infer the speaker's true intent. While performance consistently improves from the introduction of more visual cues, the overall accuracy remains far below human performance, highlighting a critical gap in multimodal reasoning. Analysis of failure cases demonstrates that current models fail to distinguish true intent from superficial correlations in the visual scene, indicating that they perceive images but do not effectively reason with them. We release our code and data at https://hazel-heejeong-nam.github.io/vague/.

多模态推理意图理解视觉上下文基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。