arXiv:2509.25183cs.CV2025-09SIGGRAPH被引 7

从随意拍摄视频中重建可变形3D物体,支持大形变和复杂运动。

PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos

  • 用个性化姿态估计器引导3D高斯表示优化,结合生成先验与可微渲染。
  • 在长视频上实现稳定追踪,重建结果在多个挑战场景中表现优异。
  • 适合动态场景理解与3D内容生成,无需特定类别先验。

我们提出PAD3R,一种从随意拍摄、无固定姿态的单目视频中重建可变形3D物体的方法。与现有方法不同,PAD3R能处理包含显著物体形变、大范围相机运动和有限视角覆盖的长视频序列,这些通常会挑战传统系统。其核心是训练一个个性化的、以物体为中心的姿态估计器,由预训练的图像到3D模型监督,指导可变形3D高斯表示的优化。优化过程通过在整个输入视频上进行长期2D点追踪进一步正则化。结合生成先验与可微渲染,PAD3R以类别无关的方式重建高保真、有结构的3D表示。大量定性和定量实验表明,PAD3R在复杂场景中鲁棒且泛化能力强,展现出其在动态场景理解与3D内容生成中的潜力。

原文摘要 · Abstract (English)

We present PAD3R, a method for reconstructing deformable 3D objects from casually captured, unposed monocular videos. Unlike existing approaches, PAD3R handles long video sequences featuring substantial object deformation, large-scale camera movement, and limited view coverage that typically challenge conventional systems. At its core, our approach trains a personalized, object-centric pose estimator, supervised by a pre-trained image-to-3D model. This guides the optimization of deformable 3D Gaussian representation. The optimization is further regularized by long-term 2D point tracking over the entire input video. By combining generative priors and differentiable rendering, PAD3R reconstructs high-fidelity, articulated 3D representations of objects in a category-agnostic way. Extensive qualitative and quantitative results show that PAD3R is robust and generalizes well across challenging scenarios, highlighting its potential for dynamic scene understanding and 3D content creation.

3D重建可变形物体视频重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。