arXiv:2410.12814cs.CV2024-10IJCAI被引 1

用生成模型找图像分类器的失效边界,直观揭示脆弱点。

Leveraging generative models to characterize the failure conditions of image classifiers

  • 在生成模型隐空间中寻找导致分类性能下降的方向
  • 发现噪声与模糊叠加的极端情况,识别出跨类与特定类的失效模式
  • 适合关注AI安全与可解释性的研究者和工程师

本文研究如何识别图像分类器的失效条件。利用近期生成对抗网络(StyleGAN2)生成高质量、可控图像数据的能力,将失效条件表达为生成模型隐空间中性能显著下降的方向。该方法能发现多种退化因素叠加的极端案例,并更细致地比较不同分类器的行为差异。通过生成图像,这些退化方向可被可视化,增强可解释性。部分退化(如图像质量)影响所有类别,而形状变化等则更具类别特异性。实验在加入噪声和模糊两种退化源的MNIST数据集上验证,展示了理解并控制关键应用中人工智能风险的潜力。

原文摘要 · Abstract (English)

We address in this work the question of identifying the failure conditions of a given image classifier. To do so, we exploit the capacity of producing controllable distributions of high quality image data made available by recent Generative Adversarial Networks (StyleGAN2): the failure conditions are expressed as directions of strong performance degradation in the generative model latent space. This strategy of analysis is used to discover corner cases that combine multiple sources of corruption, and to compare in more details the behavior of different classifiers. The directions of degradation can also be rendered visually by generating data for better interpretability. Some degradations such as image quality can affect all classes, whereas other ones such as shape are more class-specific. The approach is demonstrated on the MNIST dataset that has been completed by two sources of corruption: noise and blur, and shows a promising way to better understand and control the risks of exploiting Artificial Intelligence components for safety-critical applications.

生成模型分类器失效可解释性安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。