arXiv:2409.03605cs.CVcs.MM2024-09被引 12

用分割图分离人脸纹理与口型动作,生成更逼真的说话脸视频。

SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing

  • 以分割图为中间表示,将口型运动与纹理分离建模
  • 在HDTF和MEAD数据集上保持良好唇音同步且纹理保留更完整
  • 支持局部编辑,可替换发色、嘴唇等区域的纹理

音频驱动的说话脸生成旨在合成与输入音频同步的唇部动作视频。然而,现有生成方法在保留皮肤、牙齿等精细区域纹理方面仍存在挑战。为此,本文提出一种新框架SegTalker,通过引入分割图作为中间表示,实现唇部运动与图像纹理的解耦。具体地,利用解析网络获取的图像掩码,先由语音驱动生成动态分割图;再通过掩码引导编码器将图像语义区域分解为风格码;最后将生成的分割图与风格码注入掩码引导的StyleGAN中,生成视频帧。该方法能有效保留大部分纹理细节。此外,本方法天然支持背景分离,并可实现掩码引导的面部局部编辑:通过修改掩码并从参考图像中替换特定区域纹理(如头发、嘴唇、眉毛),实现生成过程中的无缝面部编辑。实验表明,所提方法在保持良好唇音同步的同时,显著提升了纹理保真度与视频时序一致性。在HDTF和MEAD数据集上的定量与定性结果均显示其性能优于现有方法。

原文摘要 · Abstract (English)

Audio-driven talking face generation aims to synthesize video with lip movements synchronized to input audio. However, current generative techniques face challenges in preserving intricate regional textures (skin, teeth). To address the aforementioned challenges, we propose a novel framework called SegTalker to decouple lip movements and image textures by introducing segmentation as intermediate representation. Specifically, given the mask of image employed by a parsing network, we first leverage the speech to drive the mask and generate talking segmentation. Then we disentangle semantic regions of image into style codes using a mask-guided encoder. Ultimately, we inject the previously generated talking segmentation and style codes into a mask-guided StyleGAN to synthesize video frame. In this way, most of textures are fully preserved. Moreover, our approach can inherently achieve background separation and facilitate mask-guided facial local editing. In particular, by editing the mask and swapping the region textures from a given reference image (e.g. hair, lip, eyebrows), our approach enables facial editing seamlessly when generating talking face video. Experiments demonstrate that our proposed approach can effectively preserve texture details and generate temporally consistent video while remaining competitive in lip synchronization. Quantitative and qualitative results on the HDTF and MEAD datasets illustrate the superior performance of our method over existing methods.

说话脸生成分割引导风格解耦局部编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。