用心理视觉机制构建可解释的图像表示,让模型像人一样分层理解图像。
Deep Psychovisual Image Representations

- 基于频域和复数表示学习,模仿人类视觉的中间抽象过程。
- 提取的物体局部特征更清晰可解释,优于传统CNN的模糊区域。
- 模型扩展时对深度依赖更小,适合追求效率与透明性的研究者。
心理视觉模型认为人类视觉先形成中间抽象,再进行高级认知。而深度学习模型通常使用同质的空间层堆叠提取特征,决策过程不透明。本文提出深可视编码(Deep Visual Coding),一种受1990年代图像编码启发的频率域表示,结合复数图像表示,生成类心理视觉的抽象表征。该方法利用数据驱动的频谱滤波器,在不同频段内学习任务相关的语义结构。显著性分析显示,相比传统卷积神经网络(CNN)产生的模糊区域,本模型能提取高度可解释的物体部件。此外,由于复数表示和学习到的抽象结构替代了深层空间层的作用,模型在规模扩展时对深度的依赖更小。这些发现表明,心理视觉编码为构建更高效、透明的视觉模型提供了可行路径。
原文摘要 · Abstract (English)
Psychovisual models suggest human vision decouples low-level feature extraction from higher cognition by first forming intermediate abstractions. In contrast, deep learning-based vision models routinely extract and aggregate features using homogeneous stacks of spatial layers, rendering their decision-making processes opaque. In this paper, we propose Deep Visual Coding, a learned frequency-domain representation inspired by 1990s image codes that quantised perceptually salient frequencies, which together with complex-valued image representations produces psychovisual-style abstractions. This approach enables the first psychovisual-based deep learning framework, utilizing data-driven spectral filters that learn to encode task-relevant semantic structures within distinct frequency sub-bands. Salience analyses reveal that our psychovisual models extract highly interpretable object parts compared to the amorphous regions produced by regular Convolutional Neural Networks (CNNs). Furthermore, we find that our models are less depth dependent than CNNs for model scaling, since our complex-valued representations and learned abstractions subsume the role of the deep spatial layers. Together, these findings demonstrate that psychovisual coding provides a promising path toward more efficient and transparent vision models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。