用一张到百张图生成可实时动画的逼真4D人脸,填补单图与多图重建差距。
CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models
- 基于可变形多视角扩散模型,统一处理1~100张参考图
- 单图重建质量达当前最佳,多图下逼近高精度渲染效果
- 适合影视特效、虚拟人、内容创作者快速生成动态人脸
从图像重建逼真且动态的人脸虚拟人对诸多应用至关重要,如广告、视觉特效和虚拟现实。不同应用场景对采集条件要求各异:影视工作室常使用摄像机阵列获取数百张参考图,而内容创作者可能仅需一张从网络下载的单图。目前存在大量异构的重建方法——基于多视图立体或神经渲染的技术虽质量最高,但需数百张图像;近年生成模型可从单图生成合理人脸,但视觉保真度仍落后于多视图方法。本文提出CAP4D:一种基于可变形多视角扩散模型的方法,能从任意数量(1~100张)参考图中重建逼真4D(动态3D)人脸虚拟人,并实现实时动画与渲染。该方法在单图、少图及多图场景下均达到当前最优性能,显著缩小了单图与多视图重建在视觉保真度上的差距。
原文摘要 · Abstract (English)
Reconstructing photorealistic and dynamic portrait avatars from images is essential to many applications including advertising, visual effects, and virtual reality. Depending on the application, avatar reconstruction involves different capture setups and constraints $-$ for example, visual effects studios use camera arrays to capture hundreds of reference images, while content creators may seek to animate a single portrait image downloaded from the internet. As such, there is a large and heterogeneous ecosystem of methods for avatar reconstruction. Techniques based on multi-view stereo or neural rendering achieve the highest quality results, but require hundreds of reference images. Recent generative models produce convincing avatars from a single reference image, but visual fidelity yet lags behind multi-view techniques. Here, we present CAP4D: an approach that uses a morphable multi-view diffusion model to reconstruct photoreal 4D (dynamic 3D) portrait avatars from any number of reference images (i.e., one to 100) and animate and render them in real time. Our approach demonstrates state-of-the-art performance for single-, few-, and multi-image 4D portrait avatar reconstruction, and takes steps to bridge the gap in visual fidelity between single-image and multi-view reconstruction techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。