用超像素驱动状态空间专家,让图像超分更高效精准
SP-MoMamba: Superpixel-driven Mixture of State Space Experts for Efficient Image Super-Resolution

- 以超像素为单位替代固定网格扫描,保持语义拓扑结构
- 多尺度专家动态路由,有效捕捉纹理且减少计算冗余
- 融合局部调制专家,保留高频细节,适合追求精度的场景
状态空间模型(SSMs)因其线性复杂度和长程建模能力,成为高效单图像超分辨率(SR)的有力范式。然而,现有基于Mamba的方法通常依赖数据无关的刚性扫描,将2D图像重塑为固定网格的1D序列,不可避免破坏空间-语义拓扑并引入伪影。受格式塔知觉分组理论启发,我们提出超像素驱动的状态空间专家混合模型(SP-MoMamba),实现内容感知的超分辨率。核心思想是将传统刚性扫描转化为语义级交互,以超像素为基本单元。具体地,提出超像素驱动状态空间模型(SP-SSM),将语义同质区域压缩为高阶令牌,以保持全局拓扑一致性。为解决固定扫描尺度与多样语义粒度之间的矛盾,设计多尺度超像素专家混合模块(MSS-MoE),通过动态路由机制自适应分配特定尺度专家,有效捕获多尺度纹理并降低计算冗余。此外,为防止全局抽象过程丢失高频细节,引入局部空间调制专家(LSME)补充全局建模,确保锐边与精细结构的精确重建。在标准基准上的大量实验表明,SP-MoMamba在重建保真度和效率-性能权衡上均优于当前最优的高效超分辨率方法。
原文摘要 · Abstract (English)
State space models (SSMs) have emerged as a powerful paradigm for efficient single-image super-resolution (SR) due to their linear complexity and long-range modeling capabilities. However, existing Mamba-based methods typically rely on data-agnostic rigid scanning, which reshapes 2D images into 1D sequences over a fixed grid, inevitably disrupting spatial-semantic topology and introducing artifacts. Inspired by the \textbf{Gestalt perceptual grouping theory}, we propose \textbf{SP-MoMamba}, a superpixel-driven mixture of state space experts designed for content-aware SR. Our core idea is to transform the traditional rigid scanning into a \textbf{semantic-level interaction} by treating superpixels as fundamental units. Specifically, we introduce the \textbf{Superpixel-driven State Space Model (SP-SSM)}, which compresses semantically homogeneous regions into high-order tokens to preserve global topological consistency. To address the conflict between fixed scanning scales and diverse semantic granularities, we develop the \textbf{Multi-Scale Superpixel Mixture of State Space Experts (MSS-MoE)}. This module utilizes a dynamic routing mechanism to adaptively assign scale-specific experts, effectively capturing multi-scale textures while reducing computational redundancy. Furthermore, to prevent the loss of high-frequency details during global abstraction, we introduce a \textbf{Local Spatial Modulation Expert (LSME)} to complement the global modeling, ensuring a precise reconstruction of sharp edges and fine structures. Extensive experiments on standard benchmarks demonstrate that SP-MoMamba achieves superior reconstruction fidelity and a more favorable efficiency-performance trade-off compared to state-of-the-art efficient SR methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。