仅用一张图生成完整3D人脸,支持自然动画与多视角还原。
FlexAvatar: Learning Complete 3D Head Avatars with Partial Supervision
- 引入可学习的数据源标记,统一单视角与多视角训练。
- 在单图和少样本任务中实现完整3D头像重建与真实表情动画。
- 适合需要轻量输入却追求高质量3D头像的开发者或内容创作者。
我们提出FlexAvatar,一种从单张图像生成高质量完整3D头像的方法。核心挑战在于多视角数据稀缺,且单视角训练易导致3D头像不完整。我们发现根本原因在于单视角视频学习时驱动信号与目标视角的耦合。为此,提出基于Transformer的3D肖像动画模型,引入可学习的数据源标记(称为偏置下沉),实现单视角与多视角数据的统一训练。该设计在推理时融合两者优势:单视角数据提供强泛化能力,多视角监督确保完整3D结构。此外,训练过程生成平滑的潜在头像空间,支持身份插值和任意数量输入的灵活拟合。在单视图、少样本及单视角头像生成任务中广泛评估,验证了FlexAvatar的有效性。相比现有方法在视角外推上的局限,FlexAvatar能生成完整3D头像并呈现逼真面部动画。
原文摘要 · Abstract (English)
We introduce FlexAvatar, a method for creating high-quality and complete 3D head avatars from a single image. A core challenge lies in the limited availability of multi-view data and the tendency of monocular training to yield incomplete 3D head reconstructions. We identify the root cause of this issue as the entanglement between driving signal and target viewpoint when learning from monocular videos. To address this, we propose a transformer-based 3D portrait animation model with learnable data source tokens, so-called bias sinks, which enables unified training across monocular and multi-view datasets. This design leverages the strengths of both data sources during inference: strong generalization from monocular data and full 3D completeness from multi-view supervision. Furthermore, our training procedure yields a smooth latent avatar space that facilitates identity interpolation and flexible fitting to an arbitrary number of input observations. In extensive evaluations on single-view, few-shot, and monocular avatar creation tasks, we verify the efficacy of FlexAvatar. Many existing methods struggle with view extrapolation while FlexAvatar generates complete 3D head avatars with realistic facial animations. Website: https://tobias-kirschstein.github.io/flexavatar/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。