arXiv:2605.15320cs.GRcs.CV2026-05被引 1

用几张照片秒级生成可动3D头像,支持实时动画

FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar Reconstruction

论文配图:FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar Reconstruction
图 1 · 摘自论文原文
  • 多视角融合+端到端预测,直接从像素生成可驱动头像
  • 2秒完成重建,49帧/秒动画,比顶尖方法高5.5分PSNR
  • 适合快速建模、虚拟主播、游戏角色等实时应用

传统头像重建依赖逐人优化,耗时数小时或需昂贵预处理。我们提出FFAvatar,一种通用的前馈框架,可在秒级内从少量无姿态肖像图生成高质量、可动画的3D高斯头像。通过多视角查询变换器将多源图像信息融合为统一的规范高斯表示,并直接从像素端到端预测FLAME参数实现动画,避免离线提取的开销。我们设计三阶段训练流程:(i) 在超100万身份的单目视频上进行大规模预训练,学习强泛化先验;(ii) 在360度采集的小而高质量数据集上进行多视角微调,提升几何保真度与极端视角感知;(iii) 可选个性化,在500步内适配特定身份以达最高保真。大量实验表明,FFAvatar在身份保留、几何一致性与动画保真度上树立新标准。在NeRSemble基准上,相比最先进方法LAM,PSNR提升5.5。且支持实时部署:无个性化时2秒重建,有个性化时10秒,单张NVIDIA A100 GPU下可实现49帧/秒动画。

原文摘要 · Abstract (English)

Avatar reconstruction has traditionally relied on per-subject optimization that requires hours of computation or on expensive preprocessing that limits scalability. We introduce FFAvatar, a generalizable feed-forward framework that reconstructs high-quality, animatable 3D Gaussian head avatars from few-shot unposed portrait images in seconds. FFAvatar fuses information from multiple source images into a unified canonical Gaussian representation through Multi-View Query-Former, which is animated via FLAME parameters predicted end-to-end directly from pixels, eliminating the overhead of offline FLAME extraction. We further propose a three-stage training curriculum that achieves both broad generalization and high-fidelity reconstruction: (i) scalable pretraining on extensive monocular video data with over 1M identities to learn strong generalizable priors; (ii) multi-view fine-tuning on a small but high-quality dataset of 360-degree captures to enhance geometric fidelity and extreme-view awareness; and (iii) optional personalization that adapts to specific identities for maximum fidelity within 500 optimization steps. Extensive experiments demonstrate that FFAvatar sets a new standard for identity preservation, geometric consistency, and animation fidelity. On the NeRSemble benchmark, it outperforms the state-of-the-art LAM by a substantial 5.5 PSNR gain. Furthermore, FFAvatar enables real-time deployment, reconstructing avatars in 2 seconds without personalization and 10 seconds with personalization, while supporting 49 FPS animation on a single NVIDIA A100 GPU.

3D头像高斯渲染实时生成少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。