arXiv:2603.14281cs.CV2026-03

提出解耦注意力机制,让多通道图像模型更好保留各通道特异性信息。

DC-ViT: Modulating Spatial and Channel Interactions for Multi-Channel Images

  • 用解耦自注意力分离空间与跨通道信息交互路径
  • 在三个多通道成像基准上优于现有ViT方法
  • 适合病理图像等需保留通道特性的多模态分析场景

多通道成像(MCI)的训练与评估因染色协议、传感器类型和采集设置差异导致通道配置异质,限制了固定通道编码器的应用。近期多通道Vision Transformer(MC-ViT)通过统一注意力空间联合编码所有通道的块令牌来支持灵活输入,但跨通道无约束交互会导致特征稀释,削弱对关键通道语义的保留能力。为此,本文提出解耦视觉变换器(DC-ViT),采用解耦自注意力(DSA)显式调控信息共享,将令牌更新分解为两个互补路径:建模同通道结构的空间更新,以及自适应融合跨通道信息的通道级更新。该设计缓解了信息坍塌问题,同时实现选择性跨通道交互。为进一步利用增强的通道特异性表示,引入解耦聚合(DAG),使模型可学习任务相关的通道重要性。在三个MCI基准上的大量实验表明,该方法持续优于现有MC-ViT方案。

原文摘要 · Abstract (English)

Training and evaluation in multi-channel imaging (MCI) remains challenging due to heterogeneous channel configurations arising from varying staining protocols, sensor types, and acquisition settings. This heterogeneity limits the applicability of fixed-channel encoders commonly used in general computer vision. Recent Multi-Channel Vision Transformers (MC-ViTs) address this by enabling flexible channel inputs, typically by jointly encoding patch tokens from all channels within a unified attention space. However, unrestricted token interactions across channels can lead to feature dilution, reducing the ability to preserve channel-specific semantics that are critical in MCI data. To address this, we propose Decoupled Vision Transformer (DC-ViT), which explicitly regulates information sharing using Decoupled Self-Attention (DSA), which decomposes token updates into two complementary pathways: spatial updates that model intra-channel structure, and channel-wise updates that adaptively integrate cross-channel information. This decoupling mitigates informational collapse while allowing selective inter-channel interaction. To further exploit these enhanced channel-specific representations, we introduce Decoupled Aggregation (DAG), which allows the model to learn task-specific channel importances. Extensive experiments across three MCI benchmarks demonstrate consistent improvements over existing MC-ViT approaches.

多通道图像视觉变换器解耦注意力医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。