仅用一张图生成可精确控制动作表情的逼真全身说话虚拟人
One Shot, One Talk: Whole-body Talking Avatar from a Single Image
- 用单张图像结合扩散模型生成伪视频标签,实现动作泛化
- 设计3DGS-网格混合表示并引入正则化,解决伪标签噪声问题
- 适合需要快速生成高质量虚拟人的应用,如数字人、影视制作
构建逼真且可动画化的虚拟人仍需数分钟的多视角或单视角自转视频,多数方法对动作和表情控制精度不足。为突破此限制,我们提出从单张图像构建全身说话虚拟人的新方法。针对复杂动态建模与新动作/表情泛化两大挑战,利用近期基于姿态引导的图像到视频扩散模型生成不完美视频帧作为伪标签。为应对伪视频中不一致和噪声带来的动态建模难题,提出紧密耦合的3DGS-网格混合虚拟人表示,并引入多项关键正则化以缓解由伪标签引起的不一致性。在多样化主体上的大量实验表明,该方法仅需单张图像即可生成逼真、可精确动画化且富有表现力的全身说话虚拟人。
原文摘要 · Abstract (English)
Building realistic and animatable avatars still requires minutes of multi-view or monocular self-rotating videos, and most methods lack precise control over gestures and expressions. To push this boundary, we address the challenge of constructing a whole-body talking avatar from a single image. We propose a novel pipeline that tackles two critical issues: 1) complex dynamic modeling and 2) generalization to novel gestures and expressions. To achieve seamless generalization, we leverage recent pose-guided image-to-video diffusion models to generate imperfect video frames as pseudo-labels. To overcome the dynamic modeling challenge posed by inconsistent and noisy pseudo-videos, we introduce a tightly coupled 3DGS-mesh hybrid avatar representation and apply several key regularizations to mitigate inconsistencies caused by imperfect labels. Extensive experiments on diverse subjects demonstrate that our method enables the creation of a photorealistic, precisely animatable, and expressive whole-body talking avatar from just a single image.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。