arXiv:2502.20220cs.CV2025-02ICCV被引 56

仅用几张图就能生成可动画的高保真3D人脸,无需专业设备。

Avat3r: Large Animatable Gaussian Reconstruction Model for High-fidelity 3D Head Avatars

  • 从少量图像直接重建3D头像,减少推理计算开销。
  • 利用表达码通过简单注意力机制实现头像动画,效果媲美顶尖方法。
  • 支持手机拍摄、单图甚至古董雕像等非理想输入,应用广泛。

传统创建照片级3D头像需多视角工作室采集和昂贵优化,限制其在影视特效或离线渲染外的应用。为解决此问题,我们提出Avat3r,仅需少量输入图像即可回归高质量且可动画的3D头像,显著降低推理时的计算需求。具体而言,我们将大型重建模型变为可动画,并从大规模多视角视频数据集中学习强大的3D人头先验。为提升3D重建质量,采用DUSt3R的位置图与人类基础模型Sapiens的泛化特征图。动画实现的关键发现是:仅需对表情码施加简单交叉注意力即可。训练时引入不同表情的输入图像以增强鲁棒性,使模型能处理不一致输入,如含抖动的手机拍摄或单目视频帧。我们在少输入和单输入场景下与当前最优方法对比,结果表明本方法具有竞争力。最后,我们展示了模型的广泛应用性,成功从手机拍摄、单图乃至古董雕像等跨域输入中生成3D头像。

原文摘要 · Abstract (English)

Traditionally, creating photo-realistic 3D head avatars requires a studio-level multi-view capture setup and expensive optimization during test-time, limiting the use of digital human doubles to the VFX industry or offline renderings. To address this shortcoming, we present Avat3r, which regresses a high-quality and animatable 3D head avatar from just a few input images, vastly reducing compute requirements during inference. More specifically, we make Large Reconstruction Models animatable and learn a powerful prior over 3D human heads from a large multi-view video dataset. For better 3D head reconstructions, we employ position maps from DUSt3R and generalized feature maps from the human foundation model Sapiens. To animate the 3D head, our key discovery is that simple cross-attention to an expression code is already sufficient. Finally, we increase robustness by feeding input images with different expressions to our model during training, enabling the reconstruction of 3D head avatars from inconsistent inputs, e.g., an imperfect phone capture with accidental movement, or frames from a monocular video. We compare Avat3r with current state-of-the-art methods for few-input and single-input scenarios, and find that our method has a competitive advantage in both tasks. Finally, we demonstrate the wide applicability of our proposed model, creating 3D head avatars from images of different sources, smartphone captures, single images, and even out-of-domain inputs like antique busts. Project website: https://tobias-kirschstein.github.io/avat3r/

3D头像图像重建可动画多视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。