arXiv:2607.29240cs.CVcs.AI2026-07

通过选择性校准先验,让模型更听从视觉证据而非常识误判。

When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration

论文配图:When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
图 1 · 摘自论文原文
  • 根据图像条件动态调整先验偏好,只在证据强烈时修正判断。
  • 在反事实图像上错误率下降37.2%,同时保持常识图像准确率不变。
  • 适用于多种幻觉类型,适合提升视觉语言模型可靠性。

在视觉-语言模型中,当常识先验覆盖清晰的视觉证据时,会出现常识驱动的幻觉(CDH)。例如,模型可能错误声称一个明显六指的手只有五指。我们发现这类错误具有系统性:当模型对反事实(CF)图像回答错误时,其答案往往与无图像时首选的答案一致。盲目压制先验虽可修复错误,但会破坏在匹配常识(CS)图像上的正确判断。为此,我们提出选择性先验校准(SPC),以实例相关强度从图像条件得分中减去候选级别的先验偏好估计,并仅在新得分模式强烈支持替代答案时才修正原预测。大量实验表明,SPC显著提升了对CF图像的准确性,同时基本保留了对匹配CS图像的准确率。这些改进在多种CDH类别、答案选项排列及其它冲突基准上均具泛化能力,而对无冲突基准极少产生影响。

原文摘要 · Abstract (English)

In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypical state. For example, a model may report that a visibly six-fingered hand has five fingers. We show that these errors are systematically directed: when a model answers a question about a counterfactual (CF) image incorrectly, its answer often coincides with the candidate it prefers without access to the image. Suppressing this prior indiscriminately can repair CF errors, but may also disrupt correct answers on matched commonsense (CS) images, where the same prior is helpful. We therefore propose Selective Prior Calibration (SPC), which subtracts candidate-level prior-preference estimates from image-conditioned scores with an instance-dependent strength and revises the original prediction only when the resulting score pattern strongly supports an alternative. Extensive experiments demonstrate that SPC substantially improves accuracy on CF images while largely preserving accuracy on matched CS images. Furthermore, these gains generalize across CDH categories, candidate-answer permutations, and other conflict benchmarks, while SPC rarely alters predictions on benchmarks without such conflicts.

视觉语言幻觉抑制先验校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。