arXiv:2512.04314cs.CV2025-12

分离空间与通道信息,提升多光谱图像建模效果

DisentangleFormer: Spatial-Channel Decoupling for Multi-Channel Vision

  • 将空间与通道处理解耦,独立建模结构与语义特征
  • 在多个遥感与病理数据集上达到顶尖性能,计算量降低17.8%
  • 适合多光谱、红外病理等需分离光谱信息的视觉任务

视觉变压器面临根本性限制:标准自注意力同时处理空间与通道维度,导致表示纠缠,难以独立建模结构与语义依赖。这一问题在高光谱成像中尤为突出,涵盖卫星遥感与红外病理影像,其中通道承载不同的生物物理或生化信号。本文提出DisentangleFormer,通过信息论指导的解耦设计实现稳健的多通道视觉表征。其并行架构独立处理空间-令牌与通道-令牌流,最小化冗余,包含三个核心组件:(1) 并行解耦机制,实现空间与光谱维度的去相关特征学习;(2) 压缩令牌增强器,动态融合双流特征;(3) 多尺度前馈网络,补充全局注意力以捕捉细粒度依赖。在印度松林、帕维亚大学、休斯顿及大规模大地球网遥感数据集和红外病理数据集上的实验表明,DisentangleFormer持续领先现有模型。同时,在ImageNet上保持竞争力,计算量减少17.8% FLOPs。

原文摘要 · Abstract (English)

Vision Transformers face a fundamental limitation: standard self-attention jointly processes spatial and channel dimensions, leading to entangled representations that prevent independent modeling of structural and semantic dependencies. This problem is especially pronounced in hyperspectral imaging, from satellite hyperspectral remote sensing to infrared pathology imaging, where channels capture distinct biophysical or biochemical cues. We propose DisentangleFormer, an architecture that achieves robust multi-channel vision representation through principled spatial-channel decoupling. Motivated by information-theoretic principles of decorrelated representation learning, our parallel design enables independent modeling of structural and semantic cues while minimizing redundancy between spatial and channel streams. Our design integrates three core components: (1) Parallel Disentanglement: Independently processes spatial-token and channel-token streams, enabling decorrelated feature learning across spatial and spectral dimensions, (2) Squeezed Token Enhancer: An adaptive calibration module that dynamically fuses spatial and channel streams, and (3) Multi-Scale FFN: complementing global attention with multi-scale local context to capture fine-grained structural and semantic dependencies. Extensive experiments on hyperspectral benchmarks demonstrate that DisentangleFormer achieves state-of-the-art performance, consistently outperforming existing models on Indian Pine, Pavia University, and Houston, the large-scale BigEarthNet remote sensing dataset, as well as an infrared pathology dataset. Moreover, it retains competitive accuracy on ImageNet while reducing computational cost by 17.8% in FLOPs. The code will be made publicly available upon acceptance.

多光谱视觉变压器解耦遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。