用单视频实现动态三维建模,无需重新训练。
Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video
- 融合多个预训练模型,分阶段优化动态3D重建。
- 在动态4D建模上达到顶尖性能,视觉效果优异。
- 无需微调,适合快速部署于各类视频理解任务。
本文提出一种统一方法,从日常视频中理解动态场景。大型预训练视觉基础模型(如视觉-语言、视频深度预测、运动追踪和分割模型)具备强大能力,但训练单一模型实现全面4D理解仍具挑战。我们引入Uni4D,一种多阶段优化框架,利用多个预训练模型推进动态3D建模,包括静态/动态重建、相机位姿估计和密集3D运动追踪。实验表明,Uni4D在动态4D建模任务中表现达到当前最优水平,且视觉质量显著提升。值得注意的是,Uni4D无需任何重训练或微调,充分展现了复用视觉基础模型进行4D理解的有效性。
原文摘要 · Abstract (English)
This paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth prediction, motion tracking, and segmentation models, offer promising capabilities. However, training a single model for comprehensive 4D understanding remains challenging. We introduce Uni4D, a multi-stage optimization framework that harnesses multiple pretrained models to advance dynamic 3D modeling, including static/dynamic reconstruction, camera pose estimation, and dense 3D motion tracking. Our results show state-of-the-art performance in dynamic 4D modeling with superior visual quality. Notably, Uni4D requires no retraining or fine-tuning, highlighting the effectiveness of repurposing visual foundation models for 4D understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。