arXiv:2504.17788cs.CV2025-04CVPR被引 30

构建10万条动态视频相机位姿数据集,助力真实视频生成与仿真。

Dynamic Camera Poses and Where to Find Them

  • 融合专用与通用模型筛选视频,提升标注效率。
  • 通过点追踪与动态遮罩等技术,实现更优相机位姿估计。
  • 数据集规模大且多样,适合视频生成与三维重建研究者使用。

在大规模动态互联网视频上标注相机位姿对推进真实视频生成与仿真至关重要。然而,多数互联网视频不适用于位姿估计,现有先进方法在标注动态视频时仍面临显著挑战。本文提出DynPose-100K,一个大规模动态互联网视频相机位姿标注数据集。其采集流程结合任务特定模型与通用模型进行视频筛选;在位姿估计方面,融合最新点追踪、动态掩码和运动结构(structure-from-motion)技术,优于当前最优方法。分析与实验表明,DynPose-100K在规模与多样性上均具优势,为下游应用如视频生成与三维重建开辟新可能。

原文摘要 · Abstract (English)

Annotating camera poses on dynamic Internet videos at scale is critical for advancing fields like realistic video generation and simulation. However, collecting such a dataset is difficult, as most Internet videos are unsuitable for pose estimation. Furthermore, annotating dynamic Internet videos present significant challenges even for state-of-theart methods. In this paper, we introduce DynPose-100K, a large-scale dataset of dynamic Internet videos annotated with camera poses. Our collection pipeline addresses filtering using a carefully combined set of task-specific and generalist models. For pose estimation, we combine the latest techniques of point tracking, dynamic masking, and structure-from-motion to achieve improvements over the state-of-the-art approaches. Our analysis and experiments demonstrate that DynPose-100K is both large-scale and diverse across several key attributes, opening up avenues for advancements in various downstream applications.

相机位姿视频生成数据集3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。