无需遮挡图像即可生成高保真语音驱动人脸视频,提升身份一致性。
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
- 用关键点变换让输入人脸闭嘴,避免遮挡导致的信息丢失。
- 在LRS2和HDTF数据集上,身份保留与视觉质量均优于现有方法。
- 无需参考图,可防止错误复制非语音对齐的面部元素。
语音驱动人脸生成旨在生成逼真的人脸视频,重点实现音频与唇部动作的精确同步,同时保持身份相关的视觉细节。当前最先进的方法基于图像修复技术,即遮挡输入人脸下半部分,由模型根据音频生成对齐的唇部动作。为此,这些方法还需从同一视频中随机选取未遮挡的身份参考图像以保留身份信息。然而,这种常见遮挡策略存在三大问题:(1) 输入图像信息丢失,显著影响模型对视觉质量和身份细节的保留能力;(2) 身份参考图与输入图之间的差异降低重建性能;(3) 参考图可能干扰模型,导致生成与音频不对齐的非预期元素。为解决这些问题,我们提出一种无遮挡的人脸生成方法,仍保持二维人脸编辑任务特性。我们不遮挡下半脸,而是通过两步关键点引导的方法,以无配对方式将输入图像转换为闭嘴状态。随后,将处理后的无遮挡图像与音频一同输入唇部适配模型,生成合适的唇动。因此,我们的方法既不需要遮挡输入图像,也不需要身份参考图像。我们在基准数据集LRS2和HDTF上进行了实验,并开展多种消融研究以验证贡献。
原文摘要 · Abstract (English)
Audio-Driven Talking Face Generation aims at generating realistic videos of talking faces, focusing on accurate audio-lip synchronization without deteriorating any identity-related visual details. Recent state-of-the-art methods are based on inpainting, meaning that the lower half of the input face is masked, and the model fills the masked region by generating lips aligned with the given audio. Hence, to preserve identity-related visual details from the lower half, these approaches additionally require an unmasked identity reference image randomly selected from the same video. However, this common masking strategy suffers from (1) information loss in the input faces, significantly affecting the networks' ability to preserve visual quality and identity details, (2) variation between identity reference and input image degrading reconstruction performance, and (3) the identity reference negatively impacting the model, causing unintended copying of elements unaligned with the audio. To address these issues, we propose a mask-free talking face generation approach while maintaining the 2D-based face editing task. Instead of masking the lower half, we transform the input images to have closed mouths, using a two-step landmark-based approach trained in an unpaired manner. Subsequently, we provide these edited but unmasked faces to a lip adaptation model alongside the audio to generate appropriate lip movements. Thus, our approach needs neither masked input images nor identity reference images. We conduct experiments on the benchmark LRS2 and HDTF datasets and perform various ablation studies to validate our contributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。