生成高置信度伪造样本,揭示模型对异常输入的误判风险
Network Inversion for Generating Confidently Classified Counterfeits
- 用类别标签替代输入条件进行网络反演,生成独立于真实数据的合成样本
- 所生成样本在分类器上获得高置信度,且与训练分布差异显著
- 挑战了以置信度判断数据分布的常规方法,适合关注模型安全性的研究者
在视觉分类中,生成能引发模型高置信度预测的输入,对于理解模型行为和可靠性至关重要,尤其在对抗或分布外(OOD)条件下。传统对抗方法依赖扰动现有输入以欺骗模型,但存在输入依赖性,难以同时保证高置信度和与训练数据的显著差异。本文拓展网络反演技术,提出生成自信分类伪造样本(CCCs),即不依赖特定输入、与训练分布显著不同的合成样本,仍能被模型高置信分类。通过将软向量条件替换为独热类别条件,并引入独热标签与分类器输出分布间的KL散度损失实现。实验表明,模型可对完全合成的分布外输入赋予高置信度,这挑战了基于置信度阈值的许多分布外检测方法的核心假设——即高置信意味着数据属于分布内,凸显了在安全关键应用中需采用更鲁棒的不确定性估计。
原文摘要 · Abstract (English)
In vision classification, generating inputs that elicit confident predictions is key to understanding model behavior and reliability, especially under adversarial or out-of-distribution (OOD) conditions. While traditional adversarial methods rely on perturbing existing inputs to fool a model, they are inherently input-dependent and often fail to ensure both high confidence and meaningful deviation from the training data. In this work, we extend network inversion techniques to generate Confidently Classified Counterfeits (CCCs), synthetic samples that are confidently classified by the model despite being significantly different from the training distribution and independent of any specific input. We alter inversion technique by replacing soft vector conditioning with one-hot class conditioning and introducing a Kullback-Leibler divergence loss between the one-hot label and the classifier's output distribution. CCCs offer a model-centric perspective on confidence, revealing that models can assign high confidence to entirely synthetic, out-of-distribution inputs. This challenges the core assumption behind many OOD detection techniques based on thresholding prediction confidence, which assume that high-confidence outputs imply in-distribution data, and highlights the need for more robust uncertainty estimation in safety-critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。