arXiv:2506.05937cs.LGcs.AI2025-06被引 1

让深度模型在对抗攻击下仍能准确判断自己是否不确定。

Robust Adversarial Quantification via Conflict-Aware Evidential Deep Learning

  • 通过输入变换生成多样表征,检测预测冲突来校准不确定性
  • 对异常和对抗样本的覆盖降低超55%(OOD)和90%(对抗)
  • 无需重训练,计算开销小,适合高安全场景部署

深度学习模型在高风险应用中的可靠性至关重要,分布外或对抗性输入可能导致严重后果。证据深度学习(EDL)是一种高效的不确定性量化范式,通过一次前向传播将预测建模为狄利克雷分布。然而,EDL对对抗扰动输入尤为敏感,易产生过度自信错误。冲突感知证据深度学习(C-EDL)是一种轻量级后处理不确定性量化方法,无需重新训练即可提升对抗和分布外(OOD)鲁棒性。C-EDL为每个输入生成多样且任务保持的变换,通过量化表征分歧来校准不确定性估计。其冲突感知预测调整显著提升了对分布外和对抗样本的检测能力,同时保持高分布内准确率与低计算开销。实验表明,C-EDL显著优于现有先进EDL变体及竞争基线,在多个数据集、攻击类型和不确定性度量下,对分布外数据的覆盖降低约55%,对抗数据降低约90%。

原文摘要 · Abstract (English)

Reliability of deep learning models is critical for deployment in high-stakes applications, where out-of-distribution or adversarial inputs may lead to detrimental outcomes. Evidential Deep Learning, an efficient paradigm for uncertainty quantification, models predictions as Dirichlet distributions of a single forward pass. However, EDL is particularly vulnerable to adversarially perturbed inputs, making overconfident errors. Conflict-aware Evidential Deep Learning (C-EDL) is a lightweight post-hoc uncertainty quantification approach that mitigates these issues, enhancing adversarial and OOD robustness without retraining. C-EDL generates diverse, task-preserving transformations per input and quantifies representational disagreement to calibrate uncertainty estimates when needed. C-EDL's conflict-aware prediction adjustment improves detection of OOD and adversarial inputs, maintaining high in-distribution accuracy and low computational overhead. Our experimental evaluation shows that C-EDL significantly outperforms state-of-the-art EDL variants and competitive baselines, achieving substantial reductions in coverage for OOD data (up to $\approx$55%) and adversarial data (up to $\approx$90%), across a range of datasets, attack types, and uncertainty metrics.

不确定性量化对抗鲁棒性证据学习轻量级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。