测试多模态大模型在模糊指令下的推理能力,发现其常忽略隐藏问题。
Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs
- 通过诊断套件分析模型对隐含错误的识别能力
- 六款模型在复杂场景中失败率超60%,即使具备相关能力
- 要求模型主动提问可显著提升表现,适合高风险应用
多模态大语言模型(MLLMs)越来越多地部署在开放、现实环境中,输入常为混乱、不完整或不可信。与精心设计的基准测试不同,这些场景中指令常涉及缺失对象、矛盾事实、模糊指代或不可能操作。成功不仅依赖任务执行,更取决于模型发现隐性问题的能力。本文系统分析当前MLLMs在未明确说明但需上下文推断的隐性推理场景中的表现。基于涵盖四类真实故障模式的诊断套件,评估了包括o3和GPT-4o在内的六款模型,发现它们频繁无法暴露隐藏问题,尽管具备必要感知与推理能力。显式提示表明,底层能力存在但常被用户顺从性压制。进一步证明,简单推理时干预策略如谨慎角色提示和要求澄清问题,能显著恢复性能。研究揭示了当前MLLMs在推理能力与行为合规性间的持续差距,并提出实用策略以增强其在非约束环境中的可信度。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are increasingly deployed in open-ended, real-world environments where inputs are messy, underspecified, and not always trustworthy. Unlike curated benchmarks, these settings frequently involve instructions that refer to missing objects or contradictory facts, rely on ambiguous references, or request infeasible actions. In such cases, success hinges not on task execution alone, but on a model's ability to detect when something is silently wrong. This paper presents a systematic analysis of how current MLLMs handle such implicit reasoning scenarios: cases where the flaw is not explicitly stated but must be inferred from context. Using a curated diagnostic suite spanning four categories of real-world failure modes, we evaluate six MLLMs, including o3 and GPT-4o, and find that models frequently fail to surface hidden issues, even when they possess the necessary perceptual and reasoning skills. Explicit prompting reveals that the underlying capabilities exist but are often suppressed in favor of user compliance. We further show that simple inference-time interventions, such as cautious persona prompting and, in particular, requiring a clarifying question, can dramatically recover performance. Our findings highlight a persistent gap between reasoning competence and behavioral compliance in current MLLMs and suggest practical strategies for making these models more trustworthy in underconstrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。