揭露视觉问答解释系统的不一致漏洞并提出防御方法
Adversarial Attacks on VQA-NLE: Exposing and Alleviating Inconsistencies in Visual Question Answering Explanations
- 通过扰动问题和图像生成矛盾解释,暴露系统缺陷
- 新攻击策略能诱导模型产生虚假或自相矛盾的回答
- 引入外部知识可提升解释一致性,适合关注可信AI的研究者
视觉问答中的自然语言解释(VQA-NLE)旨在通过阐明模型决策过程来增强黑箱模型的透明性。然而,我们发现现有VQA-NLE系统会产生不一致的解释,且在未真正理解上下文的情况下得出结论,暴露出推理流程或解释生成机制的缺陷。为揭示这些脆弱性,我们不仅采用现有对抗策略扰动问题,还提出一种新型策略,仅对图像进行微小改动即可诱导出矛盾或虚假输出。此外,我们引入一种基于外部知识的缓解方法,以减轻此类不一致性,从而增强模型鲁棒性。在两个标准基准和两种广泛使用的VQA-NLE模型上的大量实验表明,我们的攻击有效,知识驱动的防御具有潜力,最终揭示了当前VQA-NLE系统在安全性和可靠性方面存在的紧迫问题。
原文摘要 · Abstract (English)
Natural language explanations in visual question answering (VQA-NLE) aim to make black-box models more transparent by elucidating their decision-making processes. However, we find that existing VQA-NLE systems can produce inconsistent explanations and reach conclusions without genuinely understanding the underlying context, exposing weaknesses in either their inference pipeline or explanation-generation mechanism. To highlight these vulnerabilities, we not only leverage an existing adversarial strategy to perturb questions but also propose a novel strategy that minimally alters images to induce contradictory or spurious outputs. We further introduce a mitigation method that leverages external knowledge to alleviate these inconsistencies, thereby bolstering model robustness. Extensive evaluations on two standard benchmarks and two widely used VQA-NLE models underscore the effectiveness of our attacks and the potential of knowledge-based defenses, ultimately revealing pressing security and reliability concerns in current VQA-NLE systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。