arXiv:2506.14217cs.LGcs.AI2025-06

用三种方法综合评估模型安全,发现高准确率下仍存在推理不稳定问题。

TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift

  • 融合验证、注意力熵与解释漂移评分,统一评估模型安全性。
  • 实验显示部分高准确模型解释极不稳定,暴露深层脆弱性。
  • 适合关注模型可解释性与鲁棒性的研究人员使用。

深度神经网络通常具有高精度,但在对抗性扰动和分布外数据下的可靠性仍是一大挑战。本文提出TriGuard,一个统一的安全评估框架,结合(1)形式化鲁棒性验证,(2)注意力熵以量化显著性集中程度,以及(3)一种新颖的注意力漂移分数来衡量解释稳定性。TriGuard揭示了模型准确率与可解释性之间的关键不匹配:经过验证的模型仍可能表现出不稳定的推理过程;基于注意力的信号提供了超越对抗准确率的安全洞察。在三个数据集和五种架构上的大量实验表明,TriGuard能有效识别神经网络推理中的细微脆弱性。我们进一步证明,熵正则化训练可在不牺牲性能的前提下降低解释漂移。TriGuard推动了可解释、鲁棒模型评估的前沿发展。

原文摘要 · Abstract (English)

Deep neural networks often achieve high accuracy, but ensuring their reliability under adversarial and distributional shifts remains a pressing challenge. We propose TriGuard, a unified safety evaluation framework that combines (1) formal robustness verification, (2) attribution entropy to quantify saliency concentration, and (3) a novel Attribution Drift Score measuring explanation stability. TriGuard reveals critical mismatches between model accuracy and interpretability: verified models can still exhibit unstable reasoning, and attribution-based signals provide complementary safety insights beyond adversarial accuracy. Extensive experiments across three datasets and five architectures show how TriGuard uncovers subtle fragilities in neural reasoning. We further demonstrate that entropy-regularized training reduces explanation drift without sacrificing performance. TriGuard advances the frontier in robust, interpretable model evaluation.

模型安全可解释性鲁棒性注意力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。