量化视觉语言模型幻觉预测的抗干扰能力,找出现有模型脆弱的关键组件。
How Many Counterfactuals Does It Take? Probing VLM Hallucinations Through Circuits and Causal Effects

- 通过对比真实与反事实样本,用概率差衡量幻觉的因果影响。
- 发现仅需少量反事实样本(如10-20个)即可暴露幻觉不稳定性。
- 适合研究模型鲁棒性、幻觉机制或安全评测的开发者参考。
视觉语言模型(VLMs)常产生与视觉证据不符的幻觉预测,但现有方法缺乏对这类预测在反事实扰动下稳健性的系统理解。本文研究了幻觉输出的反事实鲁棒性样本复杂度。基于事实、反事实及激活修补运行的对数概率差异,定义因果影响度量,并用于刻画幻觉预测的稳定性。利用电路发现技术(CD-T),识别出负责这些预测的模型组件,并追踪其在反事实样本中的激活差异。通过浓度不等式和因果影响分布的方差估计,推导出可靠检测幻觉不稳定的最小反事实样本数m的实证边界。
原文摘要 · Abstract (English)
Visual Language Models (VLMs) are known to produce hallucinated predictions that are not grounded in visual evidence, yet existing approaches lack a principled understanding of how robust such predictions are under counterfactual perturbations. In this work, we study the sample complexity of counterfactual robustness for hallucinated outputs in VLMs. We define a causal influence metric based on log-probability differences between factual, counterfactual, and activation-patched runs, and use it to characterize the stability of hallucinated predictions. By leveraging circuit discovery techniques (CD-T), we identify model components responsible for these predictions and track their activation differences across counterfactual samples. We then derive empirical bounds on the minimum number of counterfactual samples m required to reliably detect instability in hallucinated outputs, using concentration inequalities and variance estimates of the causal influence distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。