arXiv:2508.10934cs.CVcs.GR2025-08被引 140

ViPE从原始视频中高效估计3D相机参数与深度,支持多种场景和镜头类型。

ViPE: Video Pose Engine for 3D Geometric Perception

  • 基于视频流联合优化相机内参、运动轨迹与密集深度图
  • 在TUM/KITTI数据集上精度比现有方法提升18%~50%,单卡3-5FPS
  • 适用于自拍、电影镜头、行车记录仪等复杂场景,适合空间智能系统开发

精确的3D几何感知是众多空间智能系统的重要前提。尽管当前先进方法依赖大规模训练数据,但从真实世界视频中获取一致且精准的3D标注仍是关键挑战。本文提出ViPE,一种轻量级通用视频处理引擎,可从无约束原始视频中高效估计相机内参、相机运动以及稠密近度量级深度图。该方法对动态自拍视频、影视镜头或行车记录仪等多种场景具有鲁棒性,支持针孔、广角及360°全景等多种相机模型。我们在多个基准上验证ViPE性能:在TUM/KITTI序列上相比现有非标定姿态估计基线提升18%/50%;单张GPU下以3-5FPS运行于标准输入分辨率。我们利用ViPE标注了大规模视频数据集,包含约10万条真实互联网视频、100万条高质量生成视频及2000条全景视频,总计约9600万帧,均附带准确相机位姿与稠密深度图。我们开源了ViPE及标注数据集,旨在加速空间智能系统的发展。

原文摘要 · Abstract (English)

Accurate 3D geometric perception is an important prerequisite for a wide range of spatial AI systems. While state-of-the-art methods depend on large-scale training data, acquiring consistent and precise 3D annotations from in-the-wild videos remains a key challenge. In this work, we introduce ViPE, a handy and versatile video processing engine designed to bridge this gap. ViPE efficiently estimates camera intrinsics, camera motion, and dense, near-metric depth maps from unconstrained raw videos. It is robust to diverse scenarios, including dynamic selfie videos, cinematic shots, or dashcams, and supports various camera models such as pinhole, wide-angle, and 360° panoramas. We have benchmarked ViPE on multiple benchmarks. Notably, it outperforms existing uncalibrated pose estimation baselines by 18%/50% on TUM/KITTI sequences, and runs at 3-5FPS on a single GPU for standard input resolutions. We use ViPE to annotate a large-scale collection of videos. This collection includes around 100K real-world internet videos, 1M high-quality AI-generated videos, and 2K panoramic videos, totaling approximately 96M frames -- all annotated with accurate camera poses and dense depth maps. We open-source ViPE and the annotated dataset with the hope of accelerating the development of spatial AI systems.

3D感知视频理解深度估计相机标定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。