arXiv:2510.26819eess.AScs.AI2025-10中稿 · TASLP

仅凭语音生成高清说话人脸,突破传统依赖图像的限制。

See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement

  • 直接从语音提取外观与动态信息,无需参考图像
  • 在三个数据集上表现优于现有方法,实现高清输出
  • 适合需要纯语音驱动人脸生成的应用场景

与依赖源图像作为外观参考、用源语音生成动作的方法不同,本文提出一种新方法,直接从语音中提取信息,解决语音到说话人脸生成中的关键挑战。首先,通过语音条件扩散模型结合统计人脸先验和样本自适应加权模块,完成高质量人脸肖像生成。随后,在语音驱动的人脸动作生成阶段,将唇动、面部表情和眼动等表达性动态嵌入扩散模型的潜在空间,并利用区域增强模块优化唇部同步。为生成高分辨率输出,集成预训练的基于Transformer的离散码本与图像渲染网络,以端到端方式提升视频帧细节。实验表明,该方法在HDTF、VoxCeleb和AVSpeech数据集上均优于现有方法。值得注意的是,这是首个仅凭单一语音输入即可生成高分辨率、高质量说话人脸视频的方法。

原文摘要 · Abstract (English)

Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in speech-to-talking face. Specifically, we first employ a speech-to-face portrait generation stage, utilizing a speech-conditioned diffusion model combined with statistical facial prior and a sample-adaptive weighting module to achieve high-quality portrait generation. In the subsequent speech-driven talking face generation stage, we embed expressive dynamics such as lip movement, facial expressions, and eye movements into the latent space of the diffusion model and further optimize lip synchronization using a region-enhancement module. To generate high-resolution outputs, we integrate a pre-trained Transformer-based discrete codebook with an image rendering network, enhancing video frame details in an end-to-end manner. Experimental results demonstrate that our method outperforms existing approaches on the HDTF, VoxCeleb, and AVSpeech datasets. Notably, this is the first method capable of generating high-resolution, high-quality talking face videos exclusively from a single speech input.

语音生成高分辨率扩散模型人脸动画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。