arXiv:2508.06715cs.CV2025-08NeurIPS被引 2

用单视频重制可变形3D动态场景,提升物理真实感

Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video

  • 通过视频回溯训练,共享运动表征连接真实与合成视频
  • 在DAVIS和PointOdyssey上实现更优几何一致性与运动质量
  • 适合需要真实动态模拟的3D内容生成研究者

随着文本到图像和图像到视频生成模型的发展,可变形3D内容生成受到越来越多关注。尽管这些模型具备丰富的外观语义先验,却难以捕捉真实4D场景合成所需的物理真实性和运动动态。相比之下,真实视频能提供难以虚构的物理几何与关节运动线索。本文提出问题:能否利用真实视频的运动先验生成物理一致的4D内容?为此,我们探索从单个视频中重制可变形3D场景,以原始视频作为监督信号纠正生成运动的伪影。提出 extbf{Restage4D},一种视频条件下的4D重制几何保持框架。方法采用视频回溯训练策略,通过共享运动表示在真实基视频与合成驱动视频间建立时间桥梁。进一步引入遮挡感知刚性损失与不可见区域回溯机制,提升复杂运动下的结构与几何一致性。在DAVIS和PointOdyssey数据集上验证,结果表明几何一致性、运动质量和3D追踪性能均获提升。本方法不仅能保留新运动下的可变形结构,还可自动修正生成模型引入的误差,揭示视频先验在4D重制任务中的潜力。源代码与训练模型将公开。

原文摘要 · Abstract (English)

Creating deformable 3D content has gained increasing attention with the rise of text-to-image and image-to-video generative models. While these models provide rich semantic priors for appearance, they struggle to capture the physical realism and motion dynamics needed for authentic 4D scene synthesis. In contrast, real-world videos can provide physically grounded geometry and articulation cues that are difficult to hallucinate. One question is raised: \textit{Can we generate physically consistent 4D content by leveraging the motion priors of the real-world video}? In this work, we explore the task of reanimating deformable 3D scenes from a single video, using the original sequence as a supervisory signal to correct artifacts from synthetic motion. We introduce \textbf{Restage4D}, a geometry-preserving pipeline for video-conditioned 4D restaging. Our approach uses a video-rewinding training strategy to temporally bridge a real base video and a synthetic driving video via a shared motion representation. We further incorporate an occlusion-aware rigidity loss and a disocclusion backtracing mechanism to improve structural and geometry consistency under challenging motion. We validate Restage4D on DAVIS and PointOdyssey, demonstrating improved geometry consistency, motion quality, and 3D tracking performance. Our method not only preserves deformable structure under novel motion, but also automatically corrects errors introduced by generative models, revealing the potential of video prior in 4D restaging task. Source code and trained models will be released.

3D重建视频生成可变形建模4D重制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。