用仿生视觉前端提升CNN鲁棒性,效果比纯训练方法更好
Explicitly Modeling Subcortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness
- 设计新型神经启发前端,融合视皮层与皮下通路模拟
- 在多种扰动下鲁棒性提升9.3%,超越基础CNN
- 适合追求生物可解释性与模型稳定性的研究者
卷积神经网络(CNN)在物体识别任务中表现优异,但在面对视觉扰动和域外图像时,仍远不如生物视觉系统稳健。已有研究证明,将标准CNN与模仿灵长类初级视觉皮层(V1)的前段模块(VOneBlock)结合,可提升整体鲁棒性。本文提出早期视觉网络(EVNets),一种新型混合CNN架构,将VOneBlock与全新设计的皮下通路模块(SubcorticalBlock)结合;该模块基于神经科学计算模型,参数化以最大化与多个实验研究中报告的皮下响应对齐。无需专门优化,该组合在多数标准V1基准上提升了对齐度,并更好地建模了非经典感受野现象。此外,EVNets展现出更强的形状偏好,在包含对抗扰动、常见损坏和域偏移的综合鲁棒性评估中,性能比基础CNN高出9.3%。最后,当与先进数据增强技术结合时,EVNets进一步提升,比单独使用数据增强提升6.2%。这表明改进架构以更贴近生物学与基于训练的方法存在互补优势。
原文摘要 · Abstract (English)
Convolutional neural networks (CNNs) trained on object recognition achieve high task performance but continue to exhibit vulnerability under a range of visual perturbations and out-of-domain images, when compared with biological vision. Prior work has demonstrated that coupling a standard CNN with a front-end (VOneBlock) that mimics the primate primary visual cortex (V1) can improve overall model robustness. Expanding on this, we introduce Early Vision Networks (EVNets), a new class of hybrid CNNs that combine the VOneBlock with a novel SubcorticalBlock, whose architecture draws from computational models in neuroscience and is parameterized to maximize alignment with subcortical responses reported across multiple experimental studies. Without being optimized to do so, the assembly of the SubcorticalBlock with the VOneBlock improved V1 alignment across most standard V1 benchmarks, and better modeled extra-classical receptive field phenomena. In addition, EVNets exhibit stronger emergent shape bias and outperform the base CNN architecture by 9.3% on an aggregate benchmark of robustness evaluations, including adversarial perturbations, common corruptions, and domain shifts. Finally, we show that EVNets can be further improved when paired with a state-of-the-art data augmentation technique, surpassing the performance of the isolated data augmentation approach by 6.2% on our robustness benchmark. This result reveals complementary benefits between changes in architecture to better mimic biology and training-based machine learning approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。