arXiv:2606.24737cs.CV2026-06

用稀疏注意力捕捉光场图像跨视角关联,高效去噪

VSANet: View-aware Sparse Attention Network for Light Field Image Denoising

论文配图:VSANet: View-aware Sparse Attention Network for Light Field Image Denoising
图 1 · 摘自论文原文
  • 将光场数据统一为时空视角令牌空间,用哈希实现稀疏注意力
  • 在多个子空间中增强关键特征,线性复杂度下实现全局交互
  • 适合需要高精度光场去噪的应用,如三维成像与虚拟现实

光场(LF)图像去噪因数据的高维结构而具有挑战性。尽管不同子孔径图像间的噪声相互独立,但场景内容在视角间存在强相关性。本文提出VSANet,一种面向光场去噪的视角感知稀疏注意力网络。具体而言,设计了视角感知稀疏注意力(VSA)模块,将4D光场特征图表示为统一的空间-角度令牌空间,并通过基于局部敏感哈希的稀疏注意力实现跨视角聚合,以线性复杂度实现全局特征交互,有效利用跨视角和空间位置的相关性。此外,设计特征精炼(FR)模块,在空间、角度及基线子空间中强化重要特征。VSA与FR模块集成于序列注意力精炼模块中,构成VSANet核心。实验表明,该方法优于现有最先进的光场去噪技术。

原文摘要 · Abstract (English)

Light field (LF) image denoising is challenging due to the high-dimensional structure of LF data. While noise is independent across sub-aperture images, scene content exhibits strong cross-view correlations. We introduce VSANet, a view-aware sparse attention network for LF denoising. Specifically, we propose a view-aware sparse attention (VSA) block that represents the 4D LF feature map as a unified spatial-angular token space and performs cross-view aggregation via locality-sensitive hashing-based sparse attention. This enables global feature interactions with linear complexity, effectively exploiting LF correlations across views and spatial locations. In addition, we design a feature refinement (FR) block to emphasize informative features in spatial, angular, and epipolar subspaces. The VSA and FR blocks are integrated within a sequential attention refinement module, forming the core of VSANet. Experiments demonstrate VSANet outperforms stateof-the-art LF denoising methods.

光场图像去噪稀疏注意力深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。