arXiv:2511.01502cs.CVcs.RO2025-11被引 3

区分运动成分提升深度与自运动联合学习的精度和鲁棒性

Discriminately Treating Motion Components Evolves Joint Depth and Ego-Motion Learning

  • 分离处理不同运动类型,利用几何规律增强约束
  • 在多个数据集上达到领先性能,尤其在复杂场景下表现优异
  • 适合做3D感知、自动驾驶等需要精准运动估计的场景

无监督学习深度与自运动是近年来3D感知的重要方向。然而,现有方法多将自运动视为辅助任务,或混合所有运动类型,或忽略与深度无关的旋转运动,限制了强几何约束的引入,导致在多样环境下可靠性不足。本文提出对运动成分进行判别式处理,利用其刚性流的几何规律来同时优化深度与自运动估计。给定连续视频帧,网络首先对齐源与目标相机的光轴和成像平面,将光流通过该对齐变换后量化偏差,分别对每个自运动分量施加几何约束,实现更精准的修正。进一步地,该对齐将联合学习重构为共轴与共面形式,使深度与各平移分量可通过闭式几何关系相互推导,引入互补约束,提升深度鲁棒性。DiMoDE框架结合上述设计,在多个公开数据集及新收集的多样化真实世界数据集上均达到当前最优表现,尤其在挑战性条件下优势显著。代码将在论文发布后开源。

原文摘要 · Abstract (English)

Unsupervised learning of depth and ego-motion, two fundamental 3D perception tasks, has made significant strides in recent years. However, most methods treat ego-motion as an auxiliary task, either mixing all motion types or excluding depth-independent rotational motions in supervision. Such designs limit the incorporation of strong geometric constraints, reducing reliability and robustness under diverse conditions. This study introduces a discriminative treatment of motion components, leveraging the geometric regularities of their respective rigid flows to benefit both depth and ego-motion estimation. Given consecutive video frames, network outputs first align the optical axes and imaging planes of the source and target cameras. Optical flows between frames are transformed through these alignments, and deviations are quantified to impose geometric constraints individually on each ego-motion component, enabling more targeted refinement. These alignments further reformulate the joint learning process into coaxial and coplanar forms, where depth and each translation component can be mutually derived through closed-form geometric relationships, introducing complementary constraints that improve depth robustness. DiMoDE, a general depth and ego-motion joint learning framework incorporating these designs, achieves state-of-the-art performance on multiple public datasets and a newly collected diverse real-world dataset, particularly under challenging conditions. Our source code will be publicly available at mias.group/DiMoDE upon publication.

深度估计自运动估计无监督学习3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。