通过追踪点的运动变化,实现单目视频的高精度相对深度估计。
Seurat: From Moving Points to Depth
- 基于2D轨迹的时空变换器,从运动中推断相对深度。
- 在TAPVid-3D上零样本泛化,真实与合成数据均表现稳定。
- 适合对视频深度感知有需求的研究者或开发者。
单目视频中的准确深度估计仍具挑战性,因单视角几何存在固有歧义,关键深度线索如视差缺失。然而,人类常通过观察物体大小和间距的变化来直观感知相对深度。受此启发,我们提出一种新方法,通过分析一组追踪的2D轨迹的空间关系及其时间演变来推断相对深度。具体地,使用现成的点追踪模型捕捉2D轨迹,再通过空间与时间变换器处理这些轨迹,直接预测随时间变化的深度。在TAPVid-3D基准测试中,该方法展现出鲁棒的零样本性能,能有效从合成数据泛化到真实世界数据集。结果表明,该方法在不同领域下均能实现时间平滑、高精度的深度预测。
原文摘要 · Abstract (English)
Accurate depth estimation from monocular videos remains challenging due to ambiguities inherent in single-view geometry, as crucial depth cues like stereopsis are absent. However, humans often perceive relative depth intuitively by observing variations in the size and spacing of objects as they move. Inspired by this, we propose a novel method that infers relative depth by examining the spatial relationships and temporal evolution of a set of tracked 2D trajectories. Specifically, we use off-the-shelf point tracking models to capture 2D trajectories. Then, our approach employs spatial and temporal transformers to process these trajectories and directly infer depth changes over time. Evaluated on the TAPVid-3D benchmark, our method demonstrates robust zero-shot performance, generalizing effectively from synthetic to real-world datasets. Results indicate that our approach achieves temporally smooth, high-accuracy depth predictions across diverse domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。