arXiv:2409.17759eess.IVcs.CV2024-09CVPR被引 7

轻量级模型LGFN提升光场图像超分辨率效果。

LGFN: Lightweight Light Field Image Super-Resolution using Local Convolution Modulation and Global Attention Feature Extraction

论文配图:LGFN: Lightweight Light Field Image Super-Resolution using Local Convolution Modulation and Global Attention Feature Extraction
图 1 · 摘自论文原文
  • 结合局部特征调制与全局注意力提取多视角信息
  • 参数仅0.45M,计算量19.33G,性能优于多数大模型
  • 适合资源受限场景下的光场图像增强应用

捕捉同一场景中不同强度和方向的光线可将三维场景信息编码为四维光场(LF)图像,广泛应用于焦距重调和深度感知。光场图像超分辨率(SR)旨在突破相机传感器性能限制,提升图像分辨率。尽管现有方法已取得良好效果,但因模型过大而难以实用。本文提出轻量级模型LGFN,融合不同视角与通道的局部与全局特征。针对同一像素在不同子孔径图像中的邻近区域具有相似结构关系,设计轻量级卷积特征提取模块(DGCE),通过特征调制更优提取局部特征;针对光场图像边界外位置存在大差异,提出高效空间注意力模块(ESAM),采用可分解大核卷积扩大感受野,并引入高效通道注意力模块(ECAM)。相比现有大型模型,本模型仅含0.45M参数、19.33G FLOPs,仍保持竞争力。大量消融实验验证有效性,在NTIRE2024光场超分辨率挑战赛第2赛道(保真度与效率)中位列第二,第1赛道(保真度)位列第七。

原文摘要 · Abstract (English)

Capturing different intensity and directions of light rays at the same scene Light field (LF) can encode the 3D scene cues into a 4D LF image which has a wide range of applications (i.e. post-capture refocusing and depth sensing). LF image super-resolution (SR) aims to improve the image resolution limited by the performance of LF camera sensor. Although existing methods have achieved promising results the practical application of these models is limited because they are not lightweight enough. In this paper we propose a lightweight model named LGFN which integrates the local and global features of different views and the features of different channels for LF image SR. Specifically owing to neighboring regions of the same pixel position in different sub-aperture images exhibit similar structural relationships we design a lightweight CNN-based feature extraction module (namely DGCE) to extract local features better through feature modulation. Meanwhile as the position beyond the boundaries in the LF image presents a large disparity we propose an efficient spatial attention module (namely ESAM) which uses decomposable large-kernel convolution to obtain an enlarged receptive field and an efficient channel attention module (namely ECAM). Compared with the existing LF image SR models with large parameter our model has a parameter of 0.45M and a FLOPs of 19.33G which has achieved a competitive effect. Extensive experiments with ablation studies demonstrate the effectiveness of our proposed method which ranked the second place in the Track 2 Fidelity & Efficiency of NTIRE2024 Light Field Super Resolution Challenge and the seventh place in the Track 1 Fidelity.

光场图像超分辨率轻量模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。