arXiv:2504.02329cs.LGcs.CV2025-04中稿 · EASE 2025被引 1

对比四种深度学习测试生成器,发现它们在漏洞发现和效率上各有优劣。

Towards Assessing Deep Learning Test Input Generators

  • 对比DeepHunter等4种生成器在漏洞发现、自然度、多样性与效率上的表现
  • 简单数据集上表现好的工具在复杂数据集上可能失效,存在性能波动
  • 为不同任务和数据集选择测试生成器提供实用指导

深度学习系统在安全关键应用中日益普及,但仍面临鲁棒性问题导致的重大失败风险。尽管已有多种测试输入生成器(TIGs)用于评估深度学习鲁棒性,但对其多维度效能的全面评估仍不足。本文对四种先进TIGs——DeepHunter、DeepFault、AdvGAN和SinVAD——在故障揭示能力、自然度、多样性和效率等方面进行了综合评估。实验采用三种预训练模型(LeNet-5、VGG16、EfficientNetB3)在不同复杂度的数据集(MNIST、CIFAR-10、ImageNet-1K)上进行验证。结果表明,各TIG在故障揭示能力、测试用例生成差异性和计算效率上存在显著权衡;且其性能随数据集复杂度变化明显:部分工具在简单数据集上表现优异但在复杂数据集上下降,而另一些则保持稳定或具备更好可扩展性。研究为根据具体目标和数据特征选择合适TIG提供了实践建议,但仍有待进一步改进以满足真实安全关键系统的需求。

原文摘要 · Abstract (English)

Deep Learning (DL) systems are increasingly deployed in safety-critical applications, yet they remain vulnerable to robustness issues that can lead to significant failures. While numerous Test Input Generators (TIGs) have been developed to evaluate DL robustness, a comprehensive assessment of their effectiveness across different dimensions is still lacking. This paper presents a comprehensive assessment of four state-of-the-art TIGs--DeepHunter, DeepFault, AdvGAN, and SinVAD--across multiple critical aspects: fault-revealing capability, naturalness, diversity, and efficiency. Our empirical study leverages three pre-trained models (LeNet-5, VGG16, and EfficientNetB3) on datasets of varying complexity (MNIST, CIFAR-10, and ImageNet-1K) to evaluate TIG performance. Our findings reveal important trade-offs in robustness revealing capability, variation in test case generation, and computational efficiency across TIGs. The results also show that TIG performance varies significantly with dataset complexity, as tools that perform well on simpler datasets may struggle with more complex ones. In contrast, others maintain steadier performance or better scalability. This paper offers practical guidance for selecting appropriate TIGs aligned with specific objectives and dataset characteristics. Nonetheless, more work is needed to address TIG limitations and advance TIGs for real-world, safety-critical systems.

深度学习测试鲁棒性评估生成器对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。