用运动图生成3D场景动态,从单图预测未来运动。
MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
- 提出像素对齐的运动图表示,融合语义与功能信息。
- 基于5万段真实视频训练扩散模型,生成合理3D轨迹。
- 适合做3D运动预测与2D视频合成的研究者使用。
本文针对从真实视频中学习具有语义和功能意义的3D运动先验这一挑战,旨在实现仅凭单张输入图像预测未来3D场景运动。我们提出一种新型的像素对齐运动图(MoMap)表示方法,可由现有生成图像模型生成,以促进高效且有效的运动预测。为学习有意义的运动分布,我们从超过50,000段真实视频中构建大规模MoMap数据库,并在此基础上训练扩散模型。所提方法不仅能合成3D中的轨迹,还提出了新的2D视频合成流程:先生成MoMap,再根据其对图像进行形变,最后完成点云渲染。实验表明,该方法生成的3D场景运动具有合理性与语义一致性。
原文摘要 · Abstract (English)
This paper addresses the challenge of learning semantically and functionally meaningful 3D motion priors from real-world videos, in order to enable prediction of future 3D scene motion from a single input image. We propose a novel pixel-aligned Motion Map (MoMap) representation for 3D scene motion, which can be generated from existing generative image models to facilitate efficient and effective motion prediction. To learn meaningful distributions over motion, we create a large-scale database of MoMaps from over 50,000 real videos and train a diffusion model on these representations. Our motion generation not only synthesizes trajectories in 3D but also suggests a new pipeline for 2D video synthesis: first generate a MoMap, then warp an image accordingly and complete the warped point-based renderings. Experimental results demonstrate that our approach generates plausible and semantically consistent 3D scene motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。