用语音驱动模型提升表情操控时的口型同步精度
Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation
- 引入语音驱动头像生成模型,精准合成口型动作
- 通过相邻帧学习策略,提升多帧生成的图像真实感
- 适合需要高保真口型同步的视频表情编辑场景
语音保持的表情操控(SPFEM)旨在改变图像和视频中的面部表情,同时保留原始口型动作。尽管已有进展,但因面部表情与口型之间的复杂关系,唇部同步仍不准确。本文利用语音驱动头像生成(AD-THG)模型在合成精确口型动作方面的优势,提出新框架THFEM,将AD-THG与SPFEM结合:由音频输入生成唇部同步的连续帧,并与表情修改后的图像融合。然而,增加帧数会降低图像真实感与表情保真度。为此,我们设计相邻帧学习策略,微调AD-THG模型以预测连续帧序列,利用邻近帧信息显著提升测试阶段图像质量。大量实验表明,该框架能有效在表情变换中保持口型一致性,充分体现了融合AD-THG与SPFEM的优势。
原文摘要 · Abstract (English)
Speech-Preserving Facial Expression Manipulation (SPFEM) is an innovative technique aimed at altering facial expressions in images and videos while retaining the original mouth movements. Despite advancements, SPFEM still struggles with accurate lip synchronization due to the complex interplay between facial expressions and mouth shapes. Capitalizing on the advanced capabilities of audio-driven talking head generation (AD-THG) models in synthesizing precise lip movements, our research introduces a novel integration of these models with SPFEM. We present a new framework, Talking Head Facial Expression Manipulation (THFEM), which utilizes AD-THG models to generate frames with accurately synchronized lip movements from audio inputs and SPFEM-altered images. However, increasing the number of frames generated by AD-THG models tends to compromise the realism and expression fidelity of the images. To counter this, we develop an adjacent frame learning strategy that finetunes AD-THG models to predict sequences of consecutive frames. This strategy enables the models to incorporate information from neighboring frames, significantly improving image quality during testing. Our extensive experimental evaluations demonstrate that this framework effectively preserves mouth shapes during expression manipulations, highlighting the substantial benefits of integrating AD-THG with SPFEM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。