arXiv:2601.07603cs.CV2026-01被引 5

用任意照片快速生成可动的3D人脸,无需专业拍摄设备。

UIKA: Fast Universal Head Avatar from Pose-Free Images

  • 通过像素级对应关系将图像颜色重投影到纹理空间,实现姿态无关建模。
  • 仅需单图或手机视频输入,在多视角和单视角设置中均优于现有方法。
  • 适合需要快速生成个性化头像的虚拟人、游戏与社交应用开发者。

我们提出UIKA,一种从任意数量的无姿态输入(包括单张图像、多视角图像和手机拍摄视频)中生成可驱动高斯人脸模型的前馈方法。不同于传统需要工作室级多视角采集和长时间优化的人类特定模型重建方式,我们从模型表示、网络设计和数据准备三方面重新思考该任务。首先,引入基于UV的头像建模策略,对每张输入图像进行像素级面部对应估计,使像素颜色可从屏幕空间重投影至UV空间,独立于相机姿态与表情变化。其次,设计可学习的UV令牌,支持在屏幕和UV层面同时应用注意力机制,通过聚合所有输入视角的UV信息,解码出规范化的高斯属性。为训练大型头像模型,我们还构建了一个大规模、身份丰富的合成训练数据集。实验表明,该方法在单目和多视角设置下均显著优于现有方法。

原文摘要 · Abstract (English)

We present UIKA, a feed-forward animatable Gaussian head model from an arbitrary number of pose-free inputs, including a single image, multi-view captures, and smartphone-captured videos. Unlike the traditional avatar method, which requires a studio-level multi-view capture system and reconstructs a human-specific model through a long-time optimization process, we rethink the task through the lenses of model representation, network design, and data preparation. First, we introduce a UV-guided avatar modeling strategy, in which each input image is associated with a pixel-wise facial correspondence estimation. Such correspondence estimation allows us to reproject each valid pixel color from screen space to UV space, which is independent of camera pose and character expression. Furthermore, we design learnable UV tokens on which the attention mechanism can be applied at both the screen and UV levels. The learned UV tokens can be decoded into canonical Gaussian attributes using aggregated UV information from all input views. To train our large avatar model, we additionally prepare a large-scale, identity-rich synthetic training dataset. Our method significantly outperforms existing approaches in both monocular and multi-view settings.

3D人脸高斯渲染快速建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。