arXiv:2608.19900cs.CV2026-08

将静态3D人像转为可控制的动态4D avatar,真实还原衣物褶皱等表面动态。

AvatarDynamizer: From Static to Dynamic Human Avatars via Generative Dynamic Textures

论文配图:AvatarDynamizer: From Static to Dynamic Human Avatars via Generative Dynamic Textures
图 1 · 摘自论文原文
  • 在纹理空间建模姿态相关的表面动态,通过条件生成实现动画
  • 支持多视角一致渲染,在有限训练数据下仍保持高视觉保真度
  • 适合需要低成本高质量动态人像的虚拟演出、数字孪生场景

对于全身虚拟人像,建模表面动态对克服恐怖谷效应、实现感知真实感至关重要。现有无特定个体的方法从单目图像、视频或文本提示中恢复静态3D人像,但其骨架驱动的动画缺乏衣物褶皱等真实表面动态;而特定个体方法虽能实现高质量渲染与真实动态,却需为每个人花费高昂成本进行多视角采集。近期通用动态人像方法难以嵌入表面动态,导致多视角一致性或动态表现力受限。为此,我们提出AvatarDynamizer,一种将现成静态3D人像转化为可控、真实且多视角一致的4D人像的生成式方法。引入新颖的纹理空间表面动态嵌入机制,将人像动态建模为条件纹理生成问题。其编码器-解码器结构将姿态依赖的动态嵌入动态纹理图中,兼容预训练视频扩散模型,并解码为3D高斯用于多视角一致渲染。由于现有数据集在规模、序列长度或运动多样性上受限,我们构建了一个大规模多视角长序列数据集,覆盖多样骨骼动作与表面动态。实验表明,该方法能有效为静态人像注入忠实的表面动态,且在视觉保真度上优于现有通用方法,尤其在动态训练数据有限时表现更优。

原文摘要 · Abstract (English)

For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism. Person-agnostic methods recover static 3D avatars from monocular images, videos, or text prompts, but their skeleton-driven animations lack realistic surface dynamics such as clothing wrinkles. In contrast, person-specific methods achieve high-quality rendering and realistic dynamics, but require expensive multi-view captures for each individual. Recent generalizable dynamic avatar methods struggle to embed surface dynamics, leading to either limited multi-view consistency or dynamic expressiveness. To this end, we propose AvatarDynamizer, a generative method that transforms an off-the-shelf static 3D avatar into a controllable, realistic, and multi-view-consistent 4D avatar. We introduce a novel texture-space surface-dynamics embedding and formulate avatar dynamics modeling as conditional texture generation. Our encoder--decoder representation embeds pose-dependent dynamics into dynamic texture maps, enabling compatibility with pre-trained video diffusion models while decoding them into 3D Gaussians for multi-view consistent rendering. Since existing datasets are limited in scale, sequence length, or motion diversity, we collect a large-scale multi-view dataset with long sequences covering diverse skeletal motions and surface dynamics. Experiments show that our method effectively animates static avatars with faithful surface dynamics and outperforms competing generalizable methods in visual fidelity, especially under limited dynamic training data.

动态人像生成模型纹理生成4D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。