用视觉错觉训练模型,提升对复杂图像的识别能力。
Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models
- 将经典几何错觉作为辅助任务融入图像分类训练
- 在复杂轮廓和细纹理场景下显著提升模型泛化性能
- 适合关注感知机制与视觉模型融合的研究者
当前深度学习模型在图像分类上表现优异,主要依赖大规模数据中的统计规律,却很少融入来自感知心理学的结构化先验。为探索感知驱动归纳偏置的潜力,我们提出将经典几何视觉错觉——人类感知中广泛研究的现象——融入标准图像分类训练流程。具体而言,我们构建了一个参数化的合成几何错觉数据集,并评估了三种多源学习策略,将错觉识别任务与ImageNet分类目标相结合。实验揭示两个关键发现:(i) 将几何错觉作为辅助监督可系统性提升模型泛化能力,尤其在涉及复杂轮廓和精细纹理的视觉挑战场景中;(ii) 即使来源于传统上被认为与自然图像识别无关的合成刺激,感知驱动的归纳偏置也能增强卷积神经网络(CNN)和基于Transformer的架构的结构敏感性。这些结果展示了感知科学与机器学习的新融合,为视觉模型设计中嵌入感知先验提供了新方向。
原文摘要 · Abstract (English)
Contemporary deep learning models have achieved impressive performance in image classification by primarily leveraging statistical regularities within large datasets, but they rarely incorporate structured insights drawn directly from perceptual psychology. To explore the potential of perceptually motivated inductive biases, we propose integrating classic geometric visual illusions well-studied phenomena from human perception into standard image-classification training pipelines. Specifically, we introduce a synthetic, parametric geometric-illusion dataset and evaluate three multi-source learning strategies that combine illusion recognition tasks with ImageNet classification objectives. Our experiments reveal two key conceptual insights: (i) incorporating geometric illusions as auxiliary supervision systematically improves generalization, especially in visually challenging cases involving intricate contours and fine textures; and (ii) perceptually driven inductive biases, even when derived from synthetic stimuli traditionally considered unrelated to natural image recognition, can enhance the structural sensitivity of both CNN and transformer-based architectures. These results demonstrate a novel integration of perceptual science and machine learning and suggest new directions for embedding perceptual priors into vision model design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。