从视频生成高保真4D动态三维模型,无需逐帧优化。
ShapeGen4D: Towards High Quality 4D Shape Generation from Videos
- 端到端视频转4D形状,用时序注意力建模动态变化。
- 在真实场景视频上提升感知保真度,减少失败案例。
- 适合做动态3D重建与数字人生成的研究者。
视频条件下的4D形状生成旨在直接从输入视频恢复随时间变化的三维几何结构和视图一致的外观。本文提出一个原生的视频到4D形状生成框架,可从视频中端到端合成单一动态3D表示。该框架基于大规模预训练3D模型引入三个关键组件:(i) 时序注意力机制,使生成过程同时依赖所有帧,并输出带时间索引的动态表示;(ii) 时序感知点采样与4D隐空间锚定,促进时空一致的几何与纹理;(iii) 跨帧噪声共享,增强时间稳定性。方法能准确捕捉非刚性运动、体积变化甚至拓扑变换,无需每帧优化。在多样化的野外视频上,相比基线方法,本方法显著提升鲁棒性与感知保真度,降低失败率。
原文摘要 · Abstract (English)
Video-conditioned 4D shape generation aims to recover time-varying 3D geometry and view-consistent appearance directly from an input video. In this work, we introduce a native video-to-4D shape generation framework that synthesizes a single dynamic 3D representation end-to-end from the video. Our framework introduces three key components based on large-scale pre-trained 3D models: (i) a temporal attention that conditions generation on all frames while producing a time-indexed dynamic representation; (ii) a time-aware point sampling and 4D latent anchoring that promote temporally consistent geometry and texture; and (iii) noise sharing across frames to enhance temporal stability. Our method accurately captures non-rigid motion, volume changes, and even topological transitions without per-frame optimization. Across diverse in-the-wild videos, our method improves robustness and perceptual fidelity and reduces failure modes compared with the baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。