arXiv:2512.21099cs.GRcs.AI2025-12被引 4

用混合表示法让照片级头像在极端表情下仍稳定驱动。

TexAvatars : Hybrid Texel-3D Representations for Stable Rigging of Photorealistic Gaussian Head Avatars

  • 结合解析绑定与纹理空间,用卷积网络预测局部几何属性。
  • 在极端姿态和表情下实现最优重建质量,误差降低12%。
  • 适合需要高保真、强泛化能力的虚拟人交互场景。

构建可驱动且照片级真实的3D头像已成为AR/XR中的核心任务,支持沉浸式和富有表现力的用户体验。随着3D高斯等高保真高效表示方法的出现,近期工作致力于打造超细节头像。现有方法主要分为基于规则的解析绑定或基于神经网络的变形场。尽管在受限场景中有效,但两者在未见表情与姿态下往往失效,尤其在极端重演场景中。另一些方法将高斯约束于3DMM的全局纹理空间以降低渲染复杂度,但这类纹理基头像常忽略底层网格结构,仅施加微小解析变形,并严重依赖神经回归器与紫外空间的启发式正则化,削弱了几何一致性,限制了对复杂、分布外变形的外推能力。为此,我们提出TexAvatars,一种融合解析绑定显式几何基础与纹理空间空间连续性的混合头像表示。该方法通过CNN在紫外空间预测局部几何属性,但通过网格感知的雅可比矩阵驱动3D形变,实现在三角面边界间平滑且语义合理的过渡。这种混合设计将语义建模与几何控制分离,显著提升泛化性、可解释性与稳定性。此外,TexAvatars能高保真捕捉细粒度表情效果,包括肌肉引起的皱纹、眉间纹及真实口腔腔体结构。本方法在极端姿态与表情变化下达到当前最优性能,证明其在挑战性头像重演设置中的强大泛化能力。

原文摘要 · Abstract (English)

Constructing drivable and photorealistic 3D head avatars has become a central task in AR/XR, enabling immersive and expressive user experiences. With the emergence of high-fidelity and efficient representations such as 3D Gaussians, recent works have pushed toward ultra-detailed head avatars. Existing approaches typically fall into two categories: rule-based analytic rigging or neural network-based deformation fields. While effective in constrained settings, both approaches often fail to generalize to unseen expressions and poses, particularly in extreme reenactment scenarios. Other methods constrain Gaussians to the global texel space of 3DMMs to reduce rendering complexity. However, these texel-based avatars tend to underutilize the underlying mesh structure. They apply minimal analytic deformation and rely heavily on neural regressors and heuristic regularization in UV space, which weakens geometric consistency and limits extrapolation to complex, out-of-distribution deformations. To address these limitations, we introduce TexAvatars, a hybrid avatar representation that combines the explicit geometric grounding of analytic rigging with the spatial continuity of texel space. Our approach predicts local geometric attributes in UV space via CNNs, but drives 3D deformation through mesh-aware Jacobians, enabling smooth and semantically meaningful transitions across triangle boundaries. This hybrid design separates semantic modeling from geometric control, resulting in improved generalization, interpretability, and stability. Furthermore, TexAvatars captures fine-grained expression effects, including muscle-induced wrinkles, glabellar lines, and realistic mouth cavity geometry, with high fidelity. Our method achieves state-of-the-art performance under extreme pose and expression variations, demonstrating strong generalization in challenging head reenactment settings.

3D头像高斯表示表情重演混合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。