仅用单张或稀疏图像生成高保真可动画3D人脸,无需相机参数或表情标签。
FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed Deformation
- 用结构化查询令牌构建通用3D表征,支持任意数量输入且不依赖相机姿态和表情标签。
- 轻量级UNet解码器结合UV位置图,实现实时精细表情动态变形,还原皱纹、露齿等细节。
- 10秒轻量微调增强极端身份特征,适合游戏、影视等需快速生成高质量角色的场景。
我们提出FlexAvatar,一种灵活的大规模重建模型,可从单张或稀疏图像中生成高保真、具备详细动态变形能力的可动画3D人脸头像,无需相机位姿或表情标签。该模型采用基于Transformer的重建架构,以结构化头部查询令牌作为基准锚点,将任意数量、无相机姿态、无表情标签的输入统一聚合为稳健的规范3D表示。为实现细节丰富的动态变形,引入一个基于UV空间位置图的轻量级UNet解码器,可在实时条件下生成表达相关的精细化形变。为更好捕捉如皱纹、露齿等罕见但关键的表情,训练阶段采用数据分布调整策略以平衡训练集中这些表达的分布。此外,仅需10秒的轻量级微调即可进一步增强极端身份下的个性细节,且不影响形变质量。大量实验表明,相比先前方法,FlexAvatar在3D一致性与动态真实感方面均表现更优,为可动画3D头像创建提供了实用解决方案。
原文摘要 · Abstract (English)
We present FlexAvatar, a flexible large reconstruction model for high-fidelity 3D head avatars with detailed dynamic deformation from single or sparse images, without requiring camera poses or expression labels. It leverages a transformer-based reconstruction model with structured head query tokens as canonical anchor to aggregate flexible input-number-agnostic, camera-pose-free and expression-free inputs into a robust canonical 3D representation. For detailed dynamic deformation, we introduce a lightweight UNet decoder conditioned on UV-space position maps, which can produce detailed expression-dependent deformations in real time. To better capture rare but critical expressions like wrinkles and bared teeth, we also adopt a data distribution adjustment strategy during training to balance the distribution of these expressions in the training set. Moreover, a lightweight 10-second refinement can further enhances identity-specific details in extreme identities without affecting deformation quality. Extensive experiments demonstrate that our FlexAvatar achieves superior 3D consistency, detailed dynamic realism compared with previous methods, providing a practical solution for animatable 3D avatar creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。