首次实现跨身份舌头动态迁移,让口型更真实自然。
TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment

- 用基础模型自举训练专用舌头分割模型,无需人工标注
- 提出空间约束扩散模型,舌头边界过渡更自然,性能超基线两倍以上
- 基于视觉语言模型评估,可大规模替代专家打分
当前人脸重演系统虽能有效转移表情与姿态,但普遍忽略舌头动态,导致说话或表情变化时口腔内部不协调。本文提出首个跨身份舌头动态迁移框架,设计基于基础模型的自举流程,无需精心标注即可构建适用于真实场景的舌头分割模型;进一步引入空间约束的隐式掩码扩散模型,通过自适应掩码扩张实现自然的嘴部边界融合。大量实验表明,该方法在所有舌头相关指标上均超越基线两倍以上。此外,提出基于视觉语言模型的评估协议,可规模化复现专家标注,验证各消融变体的感知优越性。
原文摘要 · Abstract (English)
Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tongue dynamics, leading to anatomically inconsistent mouth interiors during speech and expressive motions. We introduce the first framework for cross-identity tongue dynamics transfer in face reenactment. We propose a foundation-model-assisted bootstrapping pipeline that produces a dedicated tongue segmentation model for in-the-wild reenactment without curated annotations. We further introduce a spatially constrained latent masked diffusion model for realistic tongue synthesis, with adaptive mask dilation for seamless mouth boundary transitions. Extensive experiments demonstrate improvements of more than two times over all baselines on every tongue-specific metric. We additionally propose a VLM-based evaluation protocol that replicates expert annotation at scale, confirming perceptual superiority across all ablation variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。