诊断视觉语言模型的物体幻觉,区分是受文本先验影响还是视觉感知不足。
DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models

- 设计双维度测试:文本先验干扰与视觉证据增强,分离错误来源。
- 发现不同模型对先验压力和视觉感知的敏感度差异显著。
- 适合研究模型可靠性、幻觉机制或评估视觉语言模型的开发者。
物体层级的幻觉仍是视觉语言模型(VLMs)在二元物体存在性验证任务中的核心可靠性挑战。现有基准侧重整体准确率,却很少区分错误源于感知局限还是上下文文本先验的影响,导致失败机制模糊。我们提出 DO-Bench,一个通过结构化多模态干预进行诊断的可控基准。不采用无约束环境,而是分别探测两个互补维度:先验压制维度逐步增强上下文文本先验,同时保持视觉证据不变,以评估模型对抗先验压力的能力;感知受限维度则从全场景图像逐步过渡到局部物体裁剪图,以测量模型的感知根基强度。该配对设计可将错误归因于先验抑制、感知不足或二者交互。我们进一步定义了先验鲁棒性(PriorRobust)和感知能力(PerceptionAbility)两项诊断指标,实现行为量化。对多种开源与闭源 VLM 的评估揭示系统性差异,表明物体幻觉反映的是超越整体准确率的、机制依赖的异质性失败模式。
原文摘要 · Abstract (English)
Object level hallucination remains a central reliability challenge for vision language models (VLMs), particularly in binary object existence verification. Existing benchmarks emphasize aggregate accuracy but rarely disentangle whether errors stem from perceptual limitations or from the influence of contextual textual priors, leaving underlying failure mechanisms ambiguous. We introduce DO-Bench, a controlled diagnostic benchmark that isolates these sources through structured multimodal interventions. Rather than evaluating models in unconstrained settings, DO-Bench probes two complementary dimensions: the Prior Override dimension progressively strengthens contextual textual priors while holding visual evidence constant to assess resistance to prior pressure, and the Perception-Limited dimension incrementally enhances visual evidence from full-scene context to localized object crops to measure perceptual grounding strength. This paired design enables attribution of errors to prior suppression, perceptual insufficiency, or their interaction. We further define two diagnostic metrics, PriorRobust and PerceptionAbility, to quantify these behaviors consistently. Evaluations across diverse open- and closed-source VLMs reveal systematic differences in prior sensitivity and perceptual reliability, demonstrating that object hallucination reflects heterogeneous, mechanism dependent failure patterns beyond aggregate accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。