针对全景图超分中极区畸变问题,提出多级感知畸变的可变形网络。
Multi-level distortion-aware deformable network for omnidirectional image super-resolution
- 设计三分支可变形结构,扩大感受野以捕捉大范围畸变模式
- 在多个公开数据集上优于现有方法,峰值信噪比提升0.1~0.3dB
- 适合虚拟现实、增强现实等全景图像处理场景
随着增强现实和虚拟现实应用兴起,全景图像(ODIs)处理受到越来越多关注。全景图像超分辨率(ODISR)是提升其视觉质量的关键技术。通常将球面全景图通过等距柱状投影(ERP)映射到平面,该过程引入纬度相关的几何畸变:赤道附近畸变小,两极区域图像内容被拉伸至更广区域。然而,现有ODISR方法采样范围有限,特征提取能力不足,难以捕捉大范围畸变。为此,本文提出一种多级畸变感知可变形网络(MDDN),旨在扩展采样范围与感受野。MDDN的特征提取器包含三个并行分支:一个可变形注意力模块(对应膨胀率=1路径)及两个膨胀率为2和3的膨胀可变形卷积。该结构扩大采样范围,生成密集且全面的特征,有效捕捉ERP图像中的几何畸变。各分支提取的表征通过多级特征融合模块自适应融合。此外,为降低计算开销,采用低秩分解策略优化膨胀可变形卷积。在多个公开数据集上的大量实验表明,MDDN显著优于当前最优方法,验证了其在ODISR任务中的有效性与优越性。
原文摘要 · Abstract (English)
As augmented reality and virtual reality applications gain popularity, image processing for OmniDirectional Images (ODIs) has attracted increasing attention. OmniDirectional Image Super-Resolution (ODISR) is a promising technique for enhancing the visual quality of ODIs. Before performing super-resolution, ODIs are typically projected from a spherical surface onto a plane using EquiRectangular Projection (ERP). This projection introduces latitude-dependent geometric distortion in ERP images: distortion is minimal near the equator but becomes severe toward the poles, where image content is stretched across a wider area. However, existing ODISR methods have limited sampling ranges and feature extraction capabilities, which hinder their ability to capture distorted patterns over large areas. To address this issue, we propose a novel Multi-level Distortion-aware Deformable Network (MDDN) for ODISR, designed to expand the sampling range and receptive field. Specifically, the feature extractor in MDDN comprises three parallel branches: a deformable attention mechanism (serving as the dilation=1 path) and two dilated deformable convolutions with dilation rates of 2 and 3. This architecture expands the sampling range to include more distorted patterns across wider areas, generating dense and comprehensive features that effectively capture geometric distortions in ERP images. The representations extracted from these deformable feature extractors are adaptively fused in a multi-level feature fusion module. Furthermore, to reduce computational cost, a low-rank decomposition strategy is applied to dilated deformable convolutions. Extensive experiments on publicly available datasets demonstrate that MDDN outperforms state-of-the-art methods, underscoring its effectiveness and superiority in ODISR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。