arXiv:2411.07326cs.CV2024-11NeurIPS被引 10

将三维旋转等变性融入深度估计,提升多视角3D理解能力

$SE(3)$ Equivariant Ray Embeddings for Implicit Multi-View Depth Estimation

  • 用球谐函数实现射线位置编码,保证三维旋转等变性
  • 在真实数据集上达到当前最佳立体深度估计效果
  • 适合需要精确3D建模的机器人与视觉任务

通过将几何实体(如射线)作为输入嵌入,可有效引入归纳偏置,促进多视角学习。然而,现有方法通常缺乏等变性,而等变性对3D学习至关重要。本文探索将等变多视角学习应用于深度估计,不仅强调其在计算机视觉与机器人领域的意义,也解决先前研究的局限。多数工作忽略此设定下的等变性,或仅通过数据增强近似实现,导致不同参考系间结果不一致。为此,我们提出在Perceiver IO架构中嵌入SE(3)等变性,采用球谐函数进行位置编码以确保3D旋转等变性,并设计专用等变编码器与解码器。为验证模型,我们在立体深度估计任务上进行测试,在无需显式几何约束或大量数据增强的情况下,于真实数据集上取得当前最优性能。

原文摘要 · Abstract (English)

Incorporating inductive bias by embedding geometric entities (such as rays) as input has proven successful in multi-view learning. However, the methods adopting this technique typically lack equivariance, which is crucial for effective 3D learning. Equivariance serves as a valuable inductive prior, aiding in the generation of robust multi-view features for 3D scene understanding. In this paper, we explore the application of equivariant multi-view learning to depth estimation, not only recognizing its significance for computer vision and robotics but also addressing the limitations of previous research. Most prior studies have either overlooked equivariance in this setting or achieved only approximate equivariance through data augmentation, which often leads to inconsistencies across different reference frames. To address this issue, we propose to embed $SE(3)$ equivariance into the Perceiver IO architecture. We employ Spherical Harmonics for positional encoding to ensure 3D rotation equivariance, and develop a specialized equivariant encoder and decoder within the Perceiver IO architecture. To validate our model, we applied it to the task of stereo depth estimation, achieving state of the art results on real-world datasets without explicit geometric constraints or extensive data augmentation.

深度估计等变网络多视角学习3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。