arXiv:2508.17254cs.CVcs.AI2025-08

受生物视觉启发,新模型显著提升对错觉轮廓的感知能力。

A biological vision inspired framework for machine perception of abutting grating illusory contours

  • 模仿视觉皮层结构设计多尺度特征投影与反馈交互模块
  • 在多个测试集上对错觉轮廓识别准确率显著优于现有模型
  • 适合关注类脑智能与感知对齐的研究者

高等机器智能需与人类感知认知对齐。尽管深度神经网络(DNN)在诸多实际任务中表现卓越,但近期研究表明其无法像人类一样感知错觉轮廓中的拼接条栅(abutting grating),与人类感知模式存在偏差。为此,本文提出一种受视觉皮层电路启发的新型深度网络——错觉轮廓感知网络(ICPNet)。ICPNet采用多尺度特征投影(MFP)模块提取多层次表征,并引入特征交互注意力模块(FIAM)增强前馈与反馈特征间的交互。此外,借鉴人类感知中的形状偏见,通过边缘融合模块(EFM)执行边缘检测任务,注入形状约束以引导网络聚焦前景。我们在现有的AG-MNIST测试集以及本文构建的AG-Fashion-MNIST测试集上评估该方法。实验结果表明,ICPNet对拼接条栅错觉轮廓的敏感度显著优于当前最优模型,在多个子集上的顶1准确率均有明显提升。该工作有望推动基于DNN的模型向人类级智能迈进。

原文摘要 · Abstract (English)

Higher levels of machine intelligence demand alignment with human perception and cognition. Deep neural networks (DNN) dominated machine intelligence have demonstrated exceptional performance across various real-world tasks. Nevertheless, recent evidence suggests that DNNs fail to perceive illusory contours like the abutting grating, a discrepancy that misaligns with human perception patterns. Departing from previous works, we propose a novel deep network called illusory contour perception network (ICPNet) inspired by the circuits of the visual cortex. In ICPNet, a multi-scale feature projection (MFP) module is designed to extract multi-scale representations. To boost the interaction between feedforward and feedback features, a feature interaction attention module (FIAM) is introduced. Moreover, drawing inspiration from the shape bias observed in human perception, an edge detection task conducted via the edge fusion module (EFM) injects shape constraints that guide the network to concentrate on the foreground. We assess our method on the existing AG-MNIST test set and the AG-Fashion-MNIST test sets constructed by this work. Comprehensive experimental results reveal that ICPNet is significantly more sensitive to abutting grating illusory contours than state-of-the-art models, with notable improvements in top-1 accuracy across various subsets. This work is expected to make a step towards human-level intelligence for DNN-based models.

类脑计算错觉感知视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。