arXiv:2512.14056cs.CVcs.AI2025-12被引 1

用语音驱动的面部运动补全统一了人脸生成与编辑任务

FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling

  • 将人脸编辑与生成统一为语音条件下的面部运动补全任务
  • 支持替换、插入、删除等多种局部编辑,保持唇同步与视觉连续性
  • 提出首个专门的说话人脸编辑数据集和评估指标

说话人脸编辑与生成通常被视为独立问题。本文提出将二者视为统一范式——语音条件下的面部运动补全的子任务。我们设计了FacEDiT,一种基于流匹配训练的语音条件扩散变换器,受掩码自编码器启发,学习在周围运动和语音条件下合成被遮蔽的面部运动。该框架支持局部生成与编辑(如替换、插入、删除),并确保与未编辑区域的无缝衔接。通过引入偏置注意力和时序平滑约束,提升边界连续性和唇同步效果。为解决缺乏标准编辑基准的问题,我们构建了首个说话人脸编辑数据集FacEDiTBench,包含多样编辑类型与长度,并提出新评估指标。大量实验验证:说话人脸编辑与生成可自然涌现出语音条件运动补全中;FacEDiT在保持身份一致性的前提下,实现准确、对齐语音的面部编辑及平滑视觉过渡,并有效泛化至说话人脸生成任务。

原文摘要 · Abstract (English)

Talking face editing and face generation have often been studied as distinct problems. In this work, we propose viewing both not as separate tasks but as subtasks of a unifying formulation, speech-conditional facial motion infilling. We explore facial motion infilling as a self-supervised pretext task that also serves as a unifying formulation of dynamic talking face synthesis. To instantiate this idea, we propose FacEDiT, a speech-conditional Diffusion Transformer trained with flow matching. Inspired by masked autoencoders, FacEDiT learns to synthesize masked facial motions conditioned on surrounding motions and speech. This formulation enables both localized generation and edits, such as substitution, insertion, and deletion, while ensuring seamless transitions with unedited regions. In addition, biased attention and temporal smoothness constraints enhance boundary continuity and lip synchronization. To address the lack of a standard editing benchmark, we introduce FacEDiTBench, the first dataset for talking face editing, featuring diverse edit types and lengths, along with new evaluation metrics. Extensive experiments validate that talking face editing and generation emerge as subtasks of speech-conditional motion infilling; FacEDiT produces accurate, speech-aligned facial edits with strong identity preservation and smooth visual continuity while generalizing effectively to talking face generation.

人脸生成语音驱动视频编辑扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。