用模型自解释生成反事实测试,验证视觉语言模型是否真懂因果。
Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
- 以模型自述解释为假说,生成针对性图像修改来验证其可靠性。
- 在120个样本上发现多个模型存在显著因果不忠实现象。
- 适合安全审计、可信AI研究者使用,可输出监管可用的验证报告。
视觉语言模型(VLM)常生成看似合理但未必反映真实决策依据的自然语言解释(NLE),这种表象可信与实际因果不符的差异带来技术和治理风险。本文提出解释驱动的反事实测试(EDCT),一种全自动验证流程:给定图像-问题对,先获取模型答案和解释,再将解释解析为可检验的视觉概念,通过生成式修补生成针对性反事实图像,最后利用大模型辅助分析答案与解释的变化,计算反事实一致性得分(CCS)。在120个精选的OK-VQA样本及多个VLM上,EDCT揭示了显著的忠实性差距,并生成符合监管要求的审计证据,明确指出哪些引用概念未通过因果检验。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) often produce fluent Natural Language Explanations (NLEs) that sound convincing but may not reflect the causal factors driving predictions. This mismatch of plausibility and faithfulness poses technical and governance risks. We introduce Explanation-Driven Counterfactual Testing (EDCT), a fully automated verification procedure for a target VLM that treats the model's own explanation as a falsifiable hypothesis. Given an image-question pair, EDCT: (1) obtains the model's answer and NLE, (2) parses the NLE into testable visual concepts, (3) generates targeted counterfactual edits via generative inpainting, and (4) computes a Counterfactual Consistency Score (CCS) using LLM-assisted analysis of changes in both answers and explanations. Across 120 curated OK-VQA examples and multiple VLMs, EDCT uncovers substantial faithfulness gaps and provides regulator-aligned audit artifacts indicating when cited concepts fail causal tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。