arXiv:2412.06968cs.CV2024-12CVPR被引 14

提出新型球面变换器,实现无畸变全景感知并超越当前最佳性能。

SphereUFormer: A U-Shaped Transformer for Spherical 360 Perception

  • 基于球面局部自注意力机制,在球面域直接处理全景数据。
  • 在深度估计与语义分割任务上均超越现有最先进模型表现。
  • 适合需要高精度全景理解的自动驾驶与虚拟现实应用。

本文提出一种全新的全景360°感知方法。以往多数方法依赖等距柱状投影,虽便于使用二维操作层,但引入显著图像畸变;另一些方法尝试保持球面表示,但依赖复杂卷积核,效果不佳。本文提出基于Transformer的架构,结合创新的“球面局部自注意力”及其他球面定向模块,直接在球面域进行计算,成功实现无畸变处理,并在360°深度估计与语义分割基准测试中达到领先水平。

原文摘要 · Abstract (English)

This paper proposes a novel method for omnidirectional 360$\degree$ perception. Most common previous methods relied on equirectangular projection. This representation is easily applicable to 2D operation layers but introduces distortions into the image. Other methods attempted to remove the distortions by maintaining a sphere representation but relied on complicated convolution kernels that failed to show competitive results. In this work, we introduce a transformer-based architecture that, by incorporating a novel ``Spherical Local Self-Attention'' and other spherically-oriented modules, successfully operates in the spherical domain and outperforms the state-of-the-art in 360$\degree$ perception benchmarks for depth estimation and semantic segmentation.

球面感知视觉变换器全景重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。