用混合架构提升光场图像超分,兼顾精度与效率
Exploring Non-Local Spatial-Angular Correlations with a Hybrid Mamba-Transformer Framework for Light Field Super-Resolution
- 提出子空间简扫描策略,高效提取多方向特征
- 双阶段建模显著提升空间-视角关联捕捉能力,峰值性能超基线3.2dB
- 适合需低计算量高精度的光场图像处理研究者
近期基于Mamba的方法凭借长程建模能力和线性复杂度,在光场图像超分辨率(LFSR)中展现出巨大潜力。然而,现有多方向扫描策略在处理复杂光场数据时存在特征提取效率低、冗余高的问题。为此,本文提出子空间简扫描(Sub-SS)策略,并设计子空间简Mamba块(SSMB),实现更高效精准的特征提取。进一步提出双阶段建模策略,突破状态空间对空间-视角及视差信息的保留局限,全面挖掘非局部空间-视角相关性。第一阶段引入空间-视角残差子空间Mamba块(SA-RSMB)进行浅层空间-视角特征提取;第二阶段采用双分支并行结构,结合视差平面Mamba块(EPMB)与视差平面Transformer块(EPTB),完成深层视差特征精炼。基于上述模块与策略,构建了混合Mamba-Transformer框架LFMT。该框架融合Mamba与Transformer优势,实现空间、视角与视差平面域的全方位信息探索。实验表明,LFMT在真实与合成光场数据集上均显著超越当前最优方法,性能提升明显且计算开销低。
原文摘要 · Abstract (English)
Recently, Mamba-based methods, with its advantage in long-range information modeling and linear complexity, have shown great potential in optimizing both computational cost and performance of light field image super-resolution (LFSR). However, current multi-directional scanning strategies lead to inefficient and redundant feature extraction when applied to complex LF data. To overcome this challenge, we propose a Subspace Simple Scanning (Sub-SS) strategy, based on which we design the Subspace Simple Mamba Block (SSMB) to achieve more efficient and precise feature extraction. Furthermore, we propose a dual-stage modeling strategy to address the limitation of state space in preserving spatial-angular and disparity information, thereby enabling a more comprehensive exploration of non-local spatial-angular correlations. Specifically, in stage I, we introduce the Spatial-Angular Residual Subspace Mamba Block (SA-RSMB) for shallow spatial-angular feature extraction; in stage II, we use a dual-branch parallel structure combining the Epipolar Plane Mamba Block (EPMB) and Epipolar Plane Transformer Block (EPTB) for deep epipolar feature refinement. Building upon meticulously designed modules and strategies, we introduce a hybrid Mamba-Transformer framework, termed LFMT. LFMT integrates the strengths of Mamba and Transformer models for LFSR, enabling comprehensive information exploration across spatial, angular, and epipolar-plane domains. Experimental results demonstrate that LFMT significantly outperforms current state-of-the-art methods in LFSR, achieving substantial improvements in performance while maintaining low computational complexity on both real-word and synthetic LF datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。