同时生成人物图像与深度图,实现更真实的3D肖像动画。
Joint Learning of Depth and Appearance for Portrait Image Animation
- 基于扩散模型联合学习外观与深度信息。
- 可支持深度图转图像、图像转深度等多任务生成。
- 适用于语音驱动的逼真人脸动画,输出一致3D效果。
近年来,2D肖像动画取得显著进展,大量研究利用大型生成扩散模型中的先验知识来提升高质量图像操作能力。然而,多数方法仅输出RGB图像,而视觉与3D信息的协同生成仍缺乏探索。本文提出一种基于扩散模型的肖像生成框架,联合学习外观与深度信息。该方法采用端到端扩散范式,设计包含参考网络与通道扩展扩散主干的新架构,用于学习条件联合分布。训练完成后,框架可高效适配多种下游应用,如深度到图像、图像到深度生成、肖像重光照,以及基于音频驱动的说话头动画,且保持一致的3D输出。
原文摘要 · Abstract (English)
2D portrait animation has experienced significant advancements in recent years. Much research has utilized the prior knowledge embedded in large generative diffusion models to enhance high-quality image manipulation. However, most methods only focus on generating RGB images as output, and the co-generation of consistent visual plus 3D output remains largely under-explored. In our work, we propose to jointly learn the visual appearance and depth simultaneously in a diffusion-based portrait image generator. Our method embraces the end-to-end diffusion paradigm and introduces a new architecture suitable for learning this conditional joint distribution, consisting of a reference network and a channel-expanded diffusion backbone. Once trained, our framework can be efficiently adapted to various downstream applications, such as facial depth-to-image and image-to-depth generation, portrait relighting, and audio-driven talking head animation with consistent 3D output.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。