用滑动窗口Transformer实现无损、零参数的视频转4D网格生成。
SWiT-4D: Sliding-Window Transformer for Lossless and Parameter-Free Temporal 4D Generation
- 基于滑动窗口Transformer,融合视频时序信息重建4D网格。
- 仅需不到10秒视频微调,即实现高保真几何与稳定时序一致性。
- 适合缺乏4D数据但有图像到3D生成模型的开发者使用。
尽管4D内容生成取得进展,将单目视频转换为高质量动画3D资产并显式生成4D网格仍极具挑战。大规模自然捕获的4D网格数据集稀缺,限制了纯数据驱动的视频到4D模型从头训练。而基于丰富数据集的图像到3D生成技术提供了强大先验模型,可被利用。为此,我们提出SWiT-4D,一种滑动窗口变换器,用于无损、零参数的时序4D网格生成。SWiT-4D可无缝集成任意基于扩散Transformer(DiT)的图像到3D生成器,在保持原单图前向过程的同时,跨视频帧引入时空建模,实现任意长度视频的4D网格重建。为恢复全局位移,进一步设计适用于静态相机单目视频的优化轨迹模块。实验表明,仅用一个短于10秒的视频微调,即可实现高保真几何与稳定时序一致性,具备实际部署潜力。在域内动物园测试集及挑战性域外基准(C4D、Objaverse和真实视频)上,SWiT-4D在时序平滑性上持续优于现有基线。
原文摘要 · Abstract (English)
Despite significant progress in 4D content generation, the conversion of monocular videos into high-quality animated 3D assets with explicit 4D meshes remains considerably challenging. The scarcity of large-scale, naturally captured 4D mesh datasets further limits the ability to train generalizable video-to-4D models from scratch in a purely data-driven manner. Meanwhile, advances in image-to-3D generation, supported by extensive datasets, offer powerful prior models that can be leveraged. To better utilize these priors while minimizing reliance on 4D supervision, we introduce SWiT-4D, a Sliding-Window Transformer for lossless, parameter-free temporal 4D mesh generation. SWiT-4D integrates seamlessly with any Diffusion Transformer (DiT)-based image-to-3D generator, adding spatial-temporal modeling across video frames while preserving the original single-image forward process, enabling 4D mesh reconstruction from videos of arbitrary length. To recover global translation, we further introduce an optimization-based trajectory module tailored for static-camera monocular videos. SWiT-4D demonstrates strong data efficiency: with only a single short (<10s) video for fine-tuning, it achieves high-fidelity geometry and stable temporal consistency, indicating practical deployability under extremely limited 4D supervision. Comprehensive experiments on both in-domain zoo-test sets and challenging out-of-domain benchmarks (C4D, Objaverse, and in-the-wild videos) show that SWiT-4D consistently outperforms existing baselines in temporal smoothness. Project page: https://animotionlab.github.io/SWIT4D/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。