提出三投影融合网络,解决全景图像变形问题,提升显著目标检测精度。
TDFNet: Tri-projection Deformable Fusion Network for Panoramic Salient Object Detection

- 采用ERP、Cube Map、切线投影三路编码,融合互补投影特征
- 跨投影可变形注意力模块缓解几何失真,提升特征鲁棒性
- 适合全景视觉、机器人导航等需精准目标检测的应用场景
近年来,全景显著目标检测在机器人视觉、虚拟现实等领域展现出巨大潜力。然而,将球面场景投影到2D平面不可避免引入几何失真,严重限制现有方法性能。其中,等距柱状投影(ERP)存在极区拉伸失真,立方体图投影在面间边界产生不连续,导致特征区分度下降和几何一致性受损。为此,本文提出首个三投影可变形融合网络TDFNet,通过融合不同投影表示以缓解几何失真并提升检测性能。设计跨投影可变形注意力(CDA)模块,利用不同投影间的空间对应关系构建感知几何的采样位置,引导跨投影上下文聚合,增强对投影引起的形变的鲁棒性。同时引入纬度引导融合(LGF)模块,利用球面纬度先验构建几何置信权重,自适应平衡ERP与CMP特征;LGF还引入切线投影提供的低失真语义参考,实现跨投影特征精炼与空间对齐。基于ERP、CMP与切线投影的三分支编码结构,TDFNet同步保持全局空间连续性、局部几何细节与精细边界信息。
原文摘要 · Abstract (English)
Recent years have witnessed the growing potential of panoramic salient object detection in robotic vision, virtual reality, and related applications. However, projecting spherical scenes onto 2D planes inevitably introduces geometric distortions, which fundamentally limit the effectiveness of existing projection-based methods. Specifically, Equirectangular Projection (ERP) suffers from severe polar stretching distortions, while cube map projection introduces discontinuities across cube-face boundaries, resulting in degraded feature discriminability and compromised geometric consistency. To address these limitations, we propose TDFNet, the first Tri-projection Deformable Fusion Network for panoramic salient object detection, exploiting complementary projection representations to alleviate geometric distortions and improve detection performance.Specifically, we design a cross-projection deformable attention (CDA) module that leverages spatial correspondences between different projections to construct geometry-aware sampling locations, guiding deformable attention for cross-projection contextual aggregation and enhancing robustness against projection-induced deformations. Furthermore, we introduce a latitude-guided fusion module, which utilizes spherical latitude priors to construct geometric confidence weights for adaptively balancing ERP and CMP features. Meanwhile, LGF incorporates distortion-reduced semantic references from Tangent Projection to achieve cross-projection feature refinement and spatial alignment.By constructing a three-branch encoding architecture based on ERP, CMP, and Tangent Projection, TDFNet simultaneously preserves global spatial continuity, local geometric details, and fine-grained boundary information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。