arXiv:2409.16838cs.CVq-bio.NC2024-09被引 1

模仿视网膜和初级视觉皮层,提升CNN对噪声图像的鲁棒性

Explicitly Modeling Pre-Cortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness

  • 设计新型前级模块,模拟视网膜与初级视觉皮层的早期视觉处理
  • 在多种图像退化下,模型鲁棒性提升12.3%~18.5%,清洁图像精度略有下降
  • 适用于需应对真实世界图像噪声的视觉系统,尤其适合追求鲁棒性的场景

尽管卷积神经网络(CNN)在干净图像分类上表现优异,但在面对常见图像退化时性能显著下降,限制了其实际应用。已有研究表明,在CNN前端引入模拟灵长类初级视觉皮层(V1)特征的模块可提升整体鲁棒性。本文进一步提出两种新型生物启发式CNN架构:RetinaNet(新前级+标准后端)在各类退化下相对鲁棒性提升12.3%;EVNet(在前级后增加V1模块)则取得18.5%的相对提升。该改进在所有退化类型中均有效,且在不同后端结构上具有泛化能力。结果表明,在CNN早期层中模拟多阶段早期视觉处理,能累积提升模型鲁棒性。

原文摘要 · Abstract (English)

While convolutional neural networks (CNNs) excel at clean image classification, they struggle to classify images corrupted with different common corruptions, limiting their real-world applicability. Recent work has shown that incorporating a CNN front-end block that simulates some features of the primate primary visual cortex (V1) can improve overall model robustness. Here, we expand on this approach by introducing two novel biologically-inspired CNN model families that incorporate a new front-end block designed to simulate pre-cortical visual processing. RetinaNet, a hybrid architecture containing the novel front-end followed by a standard CNN back-end, shows a relative robustness improvement of 12.3% when compared to the standard model; and EVNet, which further adds a V1 block after the pre-cortical front-end, shows a relative gain of 18.5%. The improvement in robustness was observed for all the different corruption categories, though accompanied by a small decrease in clean image accuracy, and generalized to a different back-end architecture. These findings show that simulating multiple stages of early visual processing in CNN early layers provides cumulative benefits for model robustness.

视觉鲁棒性生物启发图像退化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。