arXiv:2605.11934cs.CV2026-05

用线性复杂度实现跨模态深度超分,提升细节还原能力

Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution

论文配图:Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
图 1 · 摘自论文原文
  • 引入交互式状态空间模型,通过局部扫描实现跨模态细粒度交互
  • 在多个数据集上达到领先性能,尤其在边缘和纹理区域表现更优
  • 适合需要高效高精度深度图重建的应用场景

引导式深度超分辨率(GDSR)利用高分辨率彩色图像指导低分辨率深度图的重建。现有方法或独立处理各模态,或依赖计算量大、复杂度为二次方的注意力机制,难以建立高效且语义交互强的联合表征。本文观察到不同模态特征在提取过程中存在语义层面的相关性,因此提出一种更灵活的方法,实现模态间的密集、语义感知深度交互。为此,我们设计了一种基于交互式状态空间模型的新框架,包含跨模态局部扫描机制,可精细捕捉RGB与深度特征间的语义关联。借助Mamba架构,该框架以线性复杂度实现全局建模。此外,引入跨模态匹配变换模块,利用双模态代表性特征进一步提升交互质量。大量实验表明,本方法在多个基准上性能优于当前最优方法。

原文摘要 · Abstract (English)

Guided depth super-resolution (GDSR) reconstructs HR depth maps from LR inputs with HR RGB guidance. Existing methods either model each modality independently or rely on computationally expensive attention mechanisms with quadratic complexity, hindering the establishment of efficient and semantically interactive joint representations. In this paper, we observe that feature maps from different modalities exhibit semantic-level correlations during feature extraction. This motivates us to develop a more flexible approach enabling dense, semantically-aware deep interactions between modalities. To this end, we propose a novel GDSR framework centered around the Interactive State Space Model. Specifically, we design a cross-modal local scanning mechanism that enables fine-grained semantic interactions between RGB and depth features. Leveraging the Mamba architecture, our framework achieves global modeling with linear complexity. Furthermore, a cross-modal matching transform module is introduced to enhance interactive modeling quality by utilizing representative features from both modalities. Extensive experiments demonstrate competitive performance against state-of-the-art methods.

深度超分跨模态状态空间模型Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。