提出新基准,测试视觉模型是否因常识误判图像内容。
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
- 设计冲突场景,让图像与常识对立,检测模型倾向
- 强模型在冲突下准确率下降超30%,显示易受先验干扰
- 适合评估视觉模型可靠性,尤其关注常识误导风险
视觉语言模型在诸多评测中表现优异,但一个基础可靠性问题仍被忽视:当视觉证据与常识冲突时,模型是遵循所见内容,还是依赖常识推断?此类情况下的典型错误表现为模型无视图像事实,输出符合常识的错误答案,称为‘常识驱动幻觉’(CDH)。为此,本文提出CDH-Bench,一个专门构建视觉证据与常识冲突情境的评测基准,涵盖计数异常、关系异常和属性异常三个维度。在二选一问答与多选题任务中评估前沿视觉语言模型,引入反事实准确率(CF-Acc)、常识准确率(CS-Acc)、反事实准确率下降(CFAD)、常识崩溃率(CCR)和相对先验依赖度(RPD)等指标。结果表明,即使强模型在视觉-常识冲突下仍易受先验影响而出现判断偏差,验证其视觉保真性不足。该基准可对模型在冲突情境下的视觉忠实性进行可控诊断。
原文摘要 · Abstract (English)
Vision-language models (VLMs) achieve strong performance on many benchmarks, yet a basic reliability question remains underexplored: when visual evidence conflicts with commonsense, do models follow what is shown or what commonsense suggests? A characteristic failure in this setting is that the model overrides visual evidence and outputs the commonsense alternative. We term this phenomenon \textbf{commonsense-driven hallucination} (CDH). To evaluate it, we introduce \textbf{CDH-Bench}, a benchmark designed to create explicit \textbf{visual evidence--commonsense conflicts}. CDH-Bench covers three dimensions: \textit{counting anomalies}, \textit{relational anomalies}, and \textit{attribute anomalies}. We evaluate frontier VLMs under \textit{binary Question Answering (QA)} and \textit{multiple-choice QA}, and report metrics including \textit{Counterfactual Accuracy} (CF-Acc), \textit{Commonsense Accuracy} (CS-Acc), \textit{Counterfactual Accuracy Drop} (CFAD), \textit{Commonsense Collapse Rate} (CCR), and \textit{Relative Prior Dependency} (RPD). Results show that even strong models remain vulnerable to prior-driven normalization under visual evidence--commonsense conflict. CDH-Bench provides a controlled diagnostic of visual fidelity under visual evidence--commonsense conflict.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。