通过频谱自适应机制提升视觉模型对形状特征的感知能力
Spectral-Adaptive Modulation Networks for Visual Perception
- 基于图谱分析统一建模卷积与自注意力的频谱特性
- 提出频谱自适应调制模块,多尺度卷积优化高频响应
- 新骨干网络在图像分类/检测/分割任务上全面超越现有模型
近期研究发现2D卷积与自注意力在频谱行为上存在差异,优化其频谱特性可提升视觉模型性能。然而,理论解释仍有限,难以说明为何2D卷积在高通滤波上更有效,以及为何大卷积核会增强形状偏好、类似自注意力。本文采用图谱分析,在统一框架下理论模拟并比较2D卷积与自注意力的频率响应。结果验证了已有经验发现,并揭示节点连接性(由窗口大小调控)是塑造频谱函数的关键因素。基于此洞察,我们提出频谱自适应调制(SPAM)混合作器,利用多尺度卷积核与频谱重缩放机制,以频谱自适应方式处理视觉特征。基于SPAM,我们构建新型视觉主干网络SPANetV2。大量实验表明,SPANetV2在ImageNet-1K分类、COCO目标检测和ADE20K语义分割等多个视觉任务上均超越当前最先进模型。
原文摘要 · Abstract (English)
Recent studies have shown that 2D convolution and self-attention exhibit distinct spectral behaviors, and optimizing their spectral properties can enhance vision model performance. However, theoretical analyses remain limited in explaining why 2D convolution is more effective in high-pass filtering than self-attention and why larger kernels favor shape bias, akin to self-attention. In this paper, we employ graph spectral analysis to theoretically simulate and compare the frequency responses of 2D convolution and self-attention within a unified framework. Our results corroborate previous empirical findings and reveal that node connectivity, modulated by window size, is a key factor in shaping spectral functions. Leveraging this insight, we introduce a \textit{spectral-adaptive modulation} (SPAM) mixer, which processes visual features in a spectral-adaptive manner using multi-scale convolutional kernels and a spectral re-scaling mechanism to refine spectral components. Based on SPAM, we develop SPANetV2 as a novel vision backbone. Extensive experiments demonstrate that SPANetV2 outperforms state-of-the-art models across multiple vision tasks, including ImageNet-1K classification, COCO object detection, and ADE20K semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。