arXiv:2410.21952cs.LG2024-10被引 5

对抗训练能同时提升模型对不确定攻击的鲁棒性。

On the Robustness of Adversarial Training Against Uncertainty Attacks

  • 利用对抗训练增强模型对扰动样本的防御能力。
  • 在CIFAR-10和ImageNet上验证了不确定性估计更可信。
  • 无需额外设计即可抵御高不确定性或低置信度攻击,适合安全敏感场景。

在学习任务中,固有的噪声导致推理不可避免地伴随一定程度的不确定性。尽管不确定性量化被广泛应用,但在安全敏感场景中,确保其可信性至关重要,下游模块需依赖可靠的不确定性估计进行决策。然而,攻击者可能诱使系统输出极高不确定性(影响可用性)或过低不确定性(错误接受可疑样本)。本文从理论与实证两方面揭示:对抗训练不仅可防御对抗样本攻击,还能在常见攻击场景下自然提升不确定性估计的可靠性,无需额外防御策略。我们在公开基准RobustBench上评估多个对抗鲁棒模型,在CIFAR-10和ImageNet数据集上验证了该结论。

原文摘要 · Abstract (English)

In learning problems, the noise inherent to the task at hand hinders the possibility to infer without a certain degree of uncertainty. Quantifying this uncertainty, regardless of its wide use, assumes high relevance for security-sensitive applications. Within these scenarios, it becomes fundamental to guarantee good (i.e., trustworthy) uncertainty measures, which downstream modules can securely employ to drive the final decision-making process. However, an attacker may be interested in forcing the system to produce either (i) highly uncertain outputs jeopardizing the system's availability or (ii) low uncertainty estimates, making the system accept uncertain samples that would instead require a careful inspection (e.g., human intervention). Therefore, it becomes fundamental to understand how to obtain robust uncertainty estimates against these kinds of attacks. In this work, we reveal both empirically and theoretically that defending against adversarial examples, i.e., carefully perturbed samples that cause misclassification, additionally guarantees a more secure, trustworthy uncertainty estimate under common attack scenarios without the need for an ad-hoc defense strategy. To support our claims, we evaluate multiple adversarial-robust models from the publicly available benchmark RobustBench on the CIFAR-10 and ImageNet datasets.

对抗训练不确定性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。