arXiv:2602.18618cs.CV2026-02

输入一张人脸图、一段语音和文字提示,生成同步说话的逼真人脸视频。

Narrating For You: Prompt-guided Audio-visual Narrating Face Generation Employing Multi-entangled Latent Space

  • 通过多纠缠潜在空间融合文本、图像和语音信息
  • 实现跨模态的时空人物特征对齐,生成自然口型与表情
  • 适合影视配音、虚拟主播等需要个性化说话人生成的场景

我们提出一种新方法,仅需一张静态人脸图像、一段语音档案和目标文本,即可合成逼真的人物说话视频。模型将提示文本、驱动图像和个体语音档案编码后,送入多纠缠潜在空间,构建音频与视频模态所需的键值对与查询。该空间负责建立跨模态的时空人物特异性特征。随后,纠缠特征分别传递至各模态解码器,生成对应音频与视频输出。

原文摘要 · Abstract (English)

We present a novel approach for generating realistic speaking and talking faces by synthesizing a person's voice and facial movements from a static image, a voice profile, and a target text. The model encodes the prompt/driving text, the driving image, and the voice profile of an individual and then combines them to pass them to the multi-entangled latent space to foster key-value pairs and queries for the audio and video modality generation pipeline. The multi-entangled latent space is responsible for establishing the spatiotemporal person-specific features between the modalities. Further, entangled features are passed to the respective decoder of each modality for output audio and video generation.

人脸生成多模态语音驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。