arXiv:2605.05748cs.AI2026-05

评估安全关键目标识别中解释方法的可靠性,发现现有方法存在误导性缺陷。

Evaluating Explainability in Safety-Critical ATR Systems: Limitations of Post-Hoc Methods and Paths Toward Robust XAI

论文配图:Evaluating Explainability in Safety-Critical ATR Systems: Limitations of Post-Hoc Methods and Paths Toward Robust XAI
图 1 · 摘自论文原文
  • 从可解释性、鲁棒性等四维度系统评估主流解释方法
  • 揭示现有方法在扰动下不稳定、易产生虚假解释等问题
  • 适合军事、航空等高风险场景的AI系统开发者参考

可解释人工智能(XAI)在安全关键型自动目标识别(ATR)系统中的部署日益重要。仅具备高预测性能不足,模型决策必须可解释、可靠且可验证。本文系统评估了ATR场景下的XAI方法,涵盖基于显著性、注意力及代理模型的范式,以及近期检测感知扩展。提出以保障为导向的评估框架,建立分类体系,并从可解释性、鲁棒性、抗操纵性、可验证性四个维度进行分析。研究发现,当前后处理解释方法存在系统性局限,如虚假解释、扰动下不稳定性,以及因视觉可信导致的过度信任。结果表明,广泛使用的XAI技术可能不足以支撑安全关键部署。最后讨论对ATR系统的启示,提出迈向更稳健、因果驱动、物理感知解释方法的方向。强调需从表面可信转向支持可靠决策与系统级保障的解释方式。

原文摘要 · Abstract (English)

Explainable Artificial Intelligence (XAI) is increasingly rec ognized as essential for deploying machine learning systems in safety critical environments. In Automatic Target Recognition (ATR), where models operate on image, video, radar, and multisensor data, high pre dictive performance alone is insufficient. Model decisions must also be interpretable, reliable, and suitable for validation. This paper presents a structured evaluation of explainability methods in the context of safety-critical ATR systems: We identify major XAI paradigms, including saliency-based, attention-based, and surrogate ap proaches, as well as recent detection-aware extensions. Based on this, we formalize explainability as an assurance-oriented assessment problem, introduce a taxonomy, and assess these methods with respect to four key dimensions: interpretability, robustness, vulnerability to manipula tion, and suitability for validation and verification. The analysis identifies systematic limitations of current post-hoc explanation methods. In par ticular, we derive critical failure modes such as spurious explanations, instability under perturbations, and overtrust induced by visually con vincing outputs. These findings indicate that widely used XAI techniques may be insufficient for safety-critical deployment. Finally, we discuss implications for ATR systems and outline directions toward more robust, causally grounded, and physically informed explain ability methods. Our results emphasize the need to move beyond visually plausible explanations toward approaches that support reliable decision making and system-level assurance.

可解释AI安全关键目标识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。