用扩散模型快速生成高精度可控4D角色动画
TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
- 通过扩散三平面重定位实现高效4D生成
- 生成时间从小时缩短至秒级,支持任意长度动画
- 适合需要高质量动态3D角色的创作者与开发者
随着3D动画需求增长,从文本描述生成高保真、可控制的4D角色仍面临挑战。现有方法存在时间与几何不一致、感知伪影、运动不规则、计算成本高及动态控制有限等问题。为此,我们提出TriDiff-4D,一种基于扩散三平面重定位的新型4D生成流程。该模型采用自回归策略,通过单次扩散过程合成每帧3D内容,利用大规模3D与动作数据显式学习结构与运动先验,实现骨骼驱动的4D生成。首先从文本生成标准3D角色与对应动作序列,再通过第二个扩散模型根据动作序列驱动角色动画,支持任意长度生成。实验表明,相比现有方法,TriDiff-4D将生成时间从小时降至秒级,无需优化过程,显著提升复杂动作生成质量,保持高保真外观与精确3D几何。
原文摘要 · Abstract (English)
With the increasing demand for 3D animation, generating high-fidelity, controllable 4D avatars from textual descriptions remains a significant challenge. Despite notable efforts in 4D generative modeling, existing methods exhibit fundamental limitations that impede their broader applicability, including temporal and geometric inconsistencies, perceptual artifacts, motion irregularities, high computational costs, and limited control over dynamics. To address these challenges, we propose TriDiff-4D, a novel 4D generative pipeline that employs diffusion-based triplane re-posing to produce high-quality, temporally coherent 4D avatars. Our model adopts an auto-regressive strategy to generate 4D sequences of arbitrary length, synthesizing each 3D frame with a single diffusion process. By explicitly learning 3D structure and motion priors from large-scale 3D and motion datasets, TriDiff-4D enables skeleton-driven 4D generation that excels in temporal consistency, motion accuracy, computational efficiency, and visual fidelity. Specifically, TriDiff-4D first generates a canonical 3D avatar and a corresponding motion sequence from a text prompt, then uses a second diffusion model to animate the avatar according to the motion sequence, supporting arbitrarily long 4D generation. Experimental results demonstrate that TriDiff-4D significantly outperforms existing methods, reducing generation time from hours to seconds by eliminating the optimization process, while substantially improving the generation of complex motions with high-fidelity appearance and accurate 3D geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。