arXiv:2603.08645cs.CVcs.GR2026-03

用检索增强提升人脸表情泛化能力,无需额外数据或模型改动。

Retrieval-Augmented Gaussian Avatars: Improving Expression Generalization

  • 训练时从大量未标注表情中检索近邻替换原表情特征。
  • 在NeRSemble上实现自驱动与跨驱动场景下表情保真度提升。
  • 无需配对数据或改架构,有效增强对表情分布偏移的鲁棒性。

无模板可动画化人脸头像可通过直接从个体捕捉数据中学习表达相关的面部形变,实现高视觉保真度,避免使用参数化人脸模板和手工设计的混合形状空间。然而,由于学习到的形变仅由单一身份的观察表达监督,这类模型存在表达覆盖范围有限的问题,当驱动动作偏离训练分布时往往表现不佳。我们提出RAF(Retrieval-Augmented Faces),一种专为从数据中学习形变的无模板头像设计的简单训练时增强方法。RAF构建一个大规模未标注表达库,在训练过程中,用该库中最近邻检索出的表达特征替换部分主体表达特征,同时仍重建原始帧。这使形变场暴露于更广泛的表达条件,促进更强的身份-表达解耦,并提升对表达分布偏移的鲁棒性,且无需配对跨身份数据、额外标注或架构变更。我们进一步分析了检索增强如何增加表达多样性,并通过用户研究验证检索质量:检索到的邻居在表达和姿态上感知更接近。在NeRSemble基准上的实验表明,RAF在自驱动和跨驱动场景下均持续提升表达保真度。

原文摘要 · Abstract (English)

Template-free animatable head avatars can achieve high visual fidelity by learning expression-dependent facial deformation directly from a subject's capture, avoiding parametric face templates and hand-designed blendshape spaces. However, since learned deformation is supervised only by the expressions observed for a single identity, these models suffer from limited expression coverage and often struggle when driven by motions that deviate from the training distribution. We introduce RAF (Retrieval-Augmented Faces), a simple training-time augmentation designed for template-free head avatars that learn deformation from data. RAF constructs a large unlabeled expression bank and, during training, replaces a subset of the subject's expression features with nearest-neighbor expressions retrieved from this bank while still reconstructing the subject's original frames. This exposes the deformation field to a broader range of expression conditions, encouraging stronger identity-expression decoupling and improving robustness to expression distribution shift without requiring paired cross-identity data, additional annotations, or architectural changes. We further analyze how retrieval augmentation increases expression diversity and validate retrieval quality with a user study showing that retrieved neighbors are perceptually closer in expression and pose. Experiments on the NeRSemble benchmark demonstrate that RAF consistently improves expression fidelity over the baseline, in both self-driving and cross-driving scenarios.

人脸生成表达泛化检索增强无模板

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。