arXiv:2601.14253cs.CV2026-01被引 14

从单视角视频生成高保真4D动态物体,关键在分离形状与运动建模。

Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis

  • 分解4D合成:先生成静态3D形状,再重建逐帧运动轨迹
  • 用参考网格学习紧凑运动隐变量,实现时序一致的完整几何还原
  • 适配不同长度视频序列,适合虚拟人、动画等4D内容生成

我们提出Motion 3-to-4,一种前馈框架,仅需单个单目视频和可选的3D参考网格,即可合成高质量4D动态物体。尽管2D、视频和3D生成近年取得显著进展,4D合成仍因训练数据有限及单目视角下几何与运动恢复的固有歧义而困难。Motion 3-to-4通过将4D合成分解为静态3D形状生成与运动重建两步解决该问题。基于规范参考网格,模型学习紧凑的运动隐表示,并预测每帧顶点轨迹以恢复完整且时序连贯的几何结构。可扩展的帧级变换器进一步增强了对不同序列长度的鲁棒性。在标准基准和新构建的具备精确真实几何的测试集上评估表明,相比先前方法,Motion 3-to-4在保真度和空间一致性方面均表现更优。项目主页见https://motion3-to-4.github.io/。

原文摘要 · Abstract (English)

We present Motion 3-to-4, a feed-forward framework for synthesising high-quality 4D dynamic objects from a single monocular video and an optional 3D reference mesh. While recent advances have significantly improved 2D, video, and 3D content generation, 4D synthesis remains difficult due to limited training data and the inherent ambiguity of recovering geometry and motion from a monocular viewpoint. Motion 3-to-4 addresses these challenges by decomposing 4D synthesis into static 3D shape generation and motion reconstruction. Using a canonical reference mesh, our model learns a compact motion latent representation and predicts per-frame vertex trajectories to recover complete, temporally coherent geometry. A scalable frame-wise transformer further enables robustness to varying sequence lengths. Evaluations on both standard benchmarks and a new dataset with accurate ground-truth geometry show that Motion 3-to-4 delivers superior fidelity and spatial consistency compared to prior work. Project page is available at https://motion3-to-4.github.io/.

4D生成单目重建运动建模动态几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。