arXiv:2607.18673cs.CV2026-07

测试视觉语言模型识别物体缺失部件的能力,发现其普遍失效且难通过工具纠正。

MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts

论文配图:MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts
图 1 · 摘自论文原文
  • 构建新基准检测模型对物体缺损的感知能力。
  • 十款主流模型在缺损检测中失败率高,即使有反证据也难纠正。
  • 现有工具辅助、推理延长等方法均无效,需底层架构改进。

视觉语言模型(VLMs)常在图像中幻觉生成不存在的物体。当物体缺失关键部分时,模型面临独特挑战,源于现实知识偏差及训练数据中此类图像稀缺。我们提出 MissingBench-Verified 基准,评估模型在无法识别物体缺失关键组件时的表现。在十款领先模型中,我们观察到持续且显著的失败率,即使外部工具证据明确反驳模型视觉判断,失败仍存在。我们进一步探究赋予模型图像处理工具(如裁剪、对比度调整)是否能实现自主检查以解决此问题。结果表明,现有缓解策略——包括工具辅助验证、自主视觉推理、延长推理时间及在更易数据集上微调——改善效果微乎其微,说明当前提示或事后修正技术无法应对此类失败模式。研究揭示了当前 VLM 在检测与监控任务中的根本局限,并强调需要从架构或训练层面入手,使模型能在面对矛盾证据时突破内在预期。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for VLMs, stemming from both real-world knowledge bias and the scarcity of such images in training data. We present MissingBench-Verified, a benchmark designed to evaluate a specific and practically relevant scenario: when vision-language models fail to recognize that an essential component of an object has been removed. Across ten leading models, we observe consistent and significant failure rates that persist even when external tool evidence explicitly contradicts the model's visual perception. We further ask whether granting models access to image processing tools (e.g., cropping, contrast adjustment) enables autonomous inspection to resolve these failures. We find that existing mitigation strategies, including tool-assisted verification, autonomous visual reasoning, longer reasoning durations, and fine-tuning on an easier dataset, provide negligible improvement, indicating that this failure mode cannot be addressed through current prompting or post-hoc correction techniques. Our findings highlight a fundamental limitation of current VLM for inspection and monitoring tasks and underscore the need for architectural or training-level interventions that enable models to override internal expectations when confronted with contradictory evidence.

视觉语言模型缺陷检测模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。