仅用单张视频重建高质量4D网格动画,无需多视角数据。
V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video
- 基于3D网格生成模型,构建端到端4D动画重建流程。
- 在多种动作类型下实现高保真几何与纹理一致性。
- 适合游戏与影视领域直接使用的4D动画资产生成。
我们提出V2M4,一种从单张单目视频直接生成可用4D网格动画的新方法。不同于依赖多视角图像与视频生成模型先验的现有方法,本方法基于原生3D网格生成模型。直接将3D网格生成模型应用于4D任务中的每一帧,易导致错误的网格姿态、外观错位及几何与纹理图的一致性问题。为此,我们设计了结构化工作流:包括相机搜索与网格重定姿、条件嵌入优化以精炼网格外观、成对网格配准以保证拓扑一致性,以及全局纹理图优化以维持纹理一致性。所提方法输出与主流图形与游戏软件兼容的高质量4D动画资产。在多种动画类型和运动幅度下的实验结果证明了该方法的泛化性与有效性。
原文摘要 · Abstract (English)
We present V2M4, a novel 4D reconstruction method that directly generates a usable 4D mesh animation asset from a single monocular video. Unlike existing approaches that rely on priors from multi-view image and video generation models, our method is based on native 3D mesh generation models. Naively applying 3D mesh generation models to generate a mesh for each frame in a 4D task can lead to issues such as incorrect mesh poses, misalignment of mesh appearance, and inconsistencies in mesh geometry and texture maps. To address these problems, we propose a structured workflow that includes camera search and mesh reposing, condition embedding optimization for mesh appearance refinement, pairwise mesh registration for topology consistency, and global texture map optimization for texture consistency. Our method outputs high-quality 4D animated assets that are compatible with mainstream graphics and game software. Experimental results across a variety of animation types and motion amplitudes demonstrate the generalization and effectiveness of our method. Project page: https://windvchen.github.io/V2M4/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。