用因果模型让AI看懂梗图背后的攻击性,透明解释判断依据。
Seeing Through VisualBERT: A Causal Adventure on Memetic Landscapes
- 构建因果模型框架,结合视觉与语义概念训练VisualBERT
- 实验证明该方法能准确识别误判原因,且比传统归因更可靠
- 适合安全敏感场景的AI可解释性研究,尤其关注隐性攻击内容
检测攻击性梗图至关重要,但主流深度神经网络模型往往缺乏透明度。现有基于输入归因的方法在处理隐性攻击性梗图和非因果归因方面存在挑战。为此,我们提出一种基于结构因果模型(SCM)的框架。在此框架中,VisualBERT基于梗图输入和因果概念联合预测类别,实现可解释性。定性评估表明,该框架能有效理解模型行为,特别是判断模型是否因正确原因做出正确判断,并揭示误分类的根本原因。定量分析评估了去混杂、对抗学习和动态路由等建模选择的重要性,并与输入归因方法进行对比。令人意外的是,输入归因方法在本框架下并不保证因果性,这引发了其在安全关键应用中的可靠性质疑。
原文摘要 · Abstract (English)
Detecting offensive memes is crucial, yet standard deep neural network systems often remain opaque. Various input attribution-based methods attempt to interpret their behavior, but they face challenges with implicitly offensive memes and non-causal attributions. To address these issues, we propose a framework based on a Structural Causal Model (SCM). In this framework, VisualBERT is trained to predict the class of an input meme based on both meme input and causal concepts, allowing for transparent interpretation. Our qualitative evaluation demonstrates the framework's effectiveness in understanding model behavior, particularly in determining whether the model was right due to the right reason, and in identifying reasons behind misclassification. Additionally, quantitative analysis assesses the significance of proposed modelling choices, such as de-confounding, adversarial learning, and dynamic routing, and compares them with input attribution methods. Surprisingly, we find that input attribution methods do not guarantee causality within our framework, raising questions about their reliability in safety-critical applications. The project page is at: https://newcodevelop.github.io/causality_adventure/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。