arXiv:2607.02516cs.CV2026-07

用对齐技术实现任意模态到视频3D的生成,无需复杂数据构建

Alignment Is All You Need For X-to-4D Generation

论文配图:Alignment Is All You Need For X-to-4D Generation
图 1 · 摘自论文原文
  • 通过物体距离与运动几何联合对齐,统一多模态输入与4D输出
  • 在X4D和Consistent4D数据集上达到最佳生成质量与一致性
  • 适合需要跨模态生成4D内容的研究者与开发者

生成式扩散模型在多模态控制下能高质量合成图像、视频和3D内容。然而,任意用户定义的模态到4D(X-to-4D)生成仍具挑战,主要因多样化数据集构建成本高且现有方法可扩展性差。本文提出Align4D框架,可将任意模态输入转化为一致的视频-3D对,利用视频引导4D运动,3D数据塑造4D几何。核心包含三项技术:(1) 物体距离对齐,分别搜索视频对齐与多视角对齐的物体距离(VAOD/MAOD),以匹配4D渲染与视频及多视角扩散先验;(2) 运动-几何联合对齐,通过同步视频与3D输入约束已知与未知视角,确保4D生成一致性;(3) 异步优化,解耦高斯属性与变形网络训练,提升运动与几何保真度。同时构建X4D数据集,整合提示、图像、视频与3D数据用于基准测试。在X4D与Consistent4D上的实验表明,Align4D在生成质量与一致性方面达到当前最优水平。

原文摘要 · Abstract (English)

Generative diffusion models excel at synthesizing high-quality images, videos, and 3D content under multimodal control. However, arbitrary user-defined modality-to-4D (X-to-4D) generation remains challenging due to the high cost of constructing diverse datasets and the limited scalability of existing methods. This paper presents Align4D, a flexible framework that translates any-modal input into coherent video-3D pairs, using video to guide 4D motion and 3D data to shape 4D geometry. Align4D introduces three key techniques: (1) Object Distance Alignment, which searches Video-Aligned and Multiview-Aligned Object Distances (VAOD/MAOD), respectively, to reconcile 4D renderings with video and the priors of multiview diffusion models; (2) Motion-Geometry Joint Alignment, which constrains known and unknown views through synchronized video and 3D inputs, ensuring consistent 4D generation; and (3) Asynchronous Optimization, which decouples Gaussian attribute and deformation network training to enhance motion and geometry fidelity. We further propose the X4D dataset, which integrates prompt, image, video, and 3D data for benchmarking. Experiments on X4D and Consistent4D demonstrate that Align4D achieves state-of-the-art quality and consistency in X-to-4D generation. Project page: https://miaoqiaowei.github.io/Align4D/.

4D生成扩散模型多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。