arXiv:2503.08636cs.LGcs.CV2025-03中稿 · Machine Learning被引 4

揭示可解释模型的欺骗风险:用原型操纵让模型误判,看似透明实则易攻。

Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning

  • 通过操控模型的潜在原型,设计对抗攻击策略。
  • 部分原型网络在攻击下准确率下降超过50%,暴露其脆弱性。
  • 适合关注AI安全与可解释性可靠性的研究者阅读。

人们普遍认为内在可解释的深度学习模型能确保正确、直观的理解,并具备更强的鲁棒性。然而,这一信念尚未得到充分验证,且越来越多证据对其提出质疑。本文揭示了这类模型因设计导致的过度依赖和易受对抗操纵的风险。我们提出了两种针对基于原型网络的对抗分析方法:原型操纵与后门攻击,并探讨了概念瓶颈模型的防御能力。通过利用模型对潜在原型的依赖,可误导其推理过程,暴露出深度神经网络本质上的不可解释性,进而引发视觉确认偏差带来的虚假安全感。部分原型网络在攻击下表现显著下降,其可信度与适用性因此受到质疑,亟需进一步研究可解释模型的鲁棒性与对齐问题。

原文摘要 · Abstract (English)

A common belief is that intrinsically interpretable deep learning models ensure a correct, intuitive understanding of their behavior and offer greater robustness against accidental errors or intentional manipulation. However, these beliefs have not been comprehensively verified, and growing evidence casts doubt on them. In this paper, we highlight the risks related to overreliance and susceptibility to adversarial manipulation of these so-called "intrinsically (aka inherently) interpretable" models by design. We introduce two strategies for adversarial analysis with prototype manipulation and backdoor attacks against prototype-based networks, and discuss how concept bottleneck models defend against these attacks. Fooling the model's reasoning by exploiting its use of latent prototypes manifests the inherent uninterpretability of deep neural networks, leading to a false sense of security reinforced by a visual confirmation bias. The reported limitations of part-prototype networks put their trustworthiness and applicability into question, motivating further work on the robustness and alignment of (deep) interpretable models.

可解释性对抗攻击原型网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。