arXiv:2508.10450cs.CV2025-08

用图像重建训练神经网络,发现其能自发产生人类感知能力。

From Images to Perception: Emergence of Perceptual Properties by Reconstructing Images

  • 用自编码、去噪等任务训练生物启发模型,不依赖人工标注感知数据。
  • 在中等噪声、模糊和稀疏条件下,模型输出与人眼判断最接近。
  • 无需监督信号,模型可自主学习符合人类感知的评价标准。

有科学家认为人类视觉感知源于图像统计特性,从而塑造早期视觉中的高效神经表征。本文提出一种生物启发架构PerceptNet,端到端优化于图像重建任务:自编码、去噪、去模糊及稀疏正则化。结果表明,编码器阶段(类V1层)虽未使用感知信息进行初始化或训练,却始终与人类对图像失真的感知判断具有最高相关性。该相关性在中等噪声、模糊和稀疏条件下达到最优。研究提示视觉系统可能针对特定失真水平与稀疏度进行优化,且生物启发模型可在无监督情况下学习感知度量。

原文摘要 · Abstract (English)

A number of scientists suggested that human visual perception may emerge from image statistics, shaping efficient neural representations in early vision. In this work, a bio-inspired architecture that can accommodate several known facts in the retina-V1 cortex, the PerceptNet, has been end-to-end optimized for different tasks related to image reconstruction: autoencoding, denoising, deblurring, and sparsity regularization. Our results show that the encoder stage (V1-like layer) consistently exhibits the highest correlation with human perceptual judgments on image distortion despite not using perceptual information in the initialization or training. This alignment exhibits an optimum for moderate noise, blur and sparsity. These findings suggest that the visual system may be tuned to remove those particular levels of distortion with that level of sparsity and that biologically inspired models can learn perceptual metrics without human supervision.

视觉感知无监督学习生物启发图像重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。