arXiv:2601.05251cs.CV2026-01被引 13

单目视频一键重建动态物体4D网格,精度更高

Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video

  • 用隐空间编码整段动画,单次前向传播完成重建
  • 结合骨骼先验训练,提升形变合理性与稳定性
  • 适合需要高效4D重建的视觉算法研究者

我们提出Mesh4D,一种基于单目视频的4D网格重建前馈模型。给定动态物体的单目视频,该模型可重建其完整3D形状与运动,以形变场形式表示。核心贡献是设计了一个紧凑的隐空间,能在一次前向传播中编码整个动画序列。该隐空间由自编码器学习,训练时受物体骨架结构引导,提供合理形变的强先验;关键在于推理阶段无需骨骼信息。编码器采用时空注意力机制,获得更稳定的整体形变表示。在此基础上,我们训练一个条件隐扩散模型,仅需输入视频和首帧网格,即可一次性预测完整动画。在重建与新视角合成基准上评估,优于现有方法,在恢复精确3D形状与形变方面表现更优。

原文摘要 · Abstract (English)

We propose Mesh4D, a feed-forward model for monocular 4D mesh reconstruction. Given a monocular video of a dynamic object, our model reconstructs the object's complete 3D shape and motion, represented as a deformation field. Our key contribution is a compact latent space that encodes the entire animation sequence in a single pass. This latent space is learned by an autoencoder that, during training, is guided by the skeletal structure of the training objects, providing strong priors on plausible deformations. Crucially, skeletal information is not required at inference time. The encoder employs spatio-temporal attention, yielding a more stable representation of the object's overall deformation. Building on this representation, we train a latent diffusion model that, conditioned on the input video and the mesh reconstructed from the first frame, predicts the full animation in one shot. We evaluate Mesh4D on reconstruction and novel view synthesis benchmarks, outperforming prior methods in recovering accurate 3D shape and deformation.

4D重建单目视频网格生成隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。