仅用两张2D图片生成连贯4D动态物体,无需模板或先验知识。
TwoSquared: 4D Generation from 2D Image Pairs
- 分两步:先从2D图生成3D结构,再用物理启发变形预测动作
- 仅需两张起止帧即可生成纹理与几何一致的4D序列
- 支持真实场景图像输入,适用于无模板通用生成
尽管生成式AI进展迅速,4D动态物体生成仍是开放挑战。受限于高质量训练数据和高计算需求,同时推测未见几何与未见运动对生成模型构成巨大难题。本文提出TwoSquared方法,仅需对应动作起始与结束的两张2D RGB图像,即可生成4D物理合理序列。该方法将问题分解为两步:1)基于高质量3D资产训练的生成模型进行图像到3D生成;2)采用物理启发的形变模块预测中间动作。本方法不依赖模板或物体类别特有先验知识,可接受真实场景图像输入。实验表明,TwoSquared仅凭两张2D图像即可生成纹理一致、几何一致的4D序列。
原文摘要 · Abstract (English)
Despite the astonishing progress in generative AI, 4D dynamic object generation remains an open challenge. With limited high-quality training data and heavy computing requirements, the combination of hallucinating unseen geometry together with unseen movement poses great challenges to generative models. In this work, we propose TwoSquared as a method to obtain a 4D physically plausible sequence starting from only two 2D RGB images corresponding to the beginning and end of the action. Instead of directly solving the 4D generation problem, TwoSquared decomposes the problem into two steps: 1) an image-to-3D module generation based on the existing generative model trained on high-quality 3D assets, and 2) a physically inspired deformation module to predict intermediate movements. To this end, our method does not require templates or object-class-specific prior knowledge and can take in-the-wild images as input. In our experiments, we demonstrate that TwoSquared is capable of producing texture-consistent and geometry-consistent 4D sequences only given 2D images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。