仅用视频生成真实语音,让无声画面开口说话。
Shushing! Let's Imagine an Authentic Speech from the Silent Video
- 用离散扩散模型从唇动预测语音,提升语义一致性。
- 结合BERT修正错误音素,生成语音更准确且情感丰富。
- 适合影视配音、语言障碍辅助等真实场景应用。
视觉引导语音生成旨在仅凭面部表情或唇部动作生成真实语音,无需音频信号,对电影配音和失语症患者辅助具有重要价值。现有方法难以在语义、音色和情感韵律上实现跨模态统一对齐,为此我们提出一致视频到语音(CV2S)任务以增强跨模态一致性。为此,我们提出ImaginTalk,一种新颖的跨模态扩散框架,仅使用视觉输入生成忠实语音,运行于离散空间。具体地,我们设计离散唇部对齐器,从唇部视频预测离散语音标记以捕捉语义信息,并通过误差检测器识别错位标记,再通过BERT进行掩码语言建模修正。为增强语音表现力,我们构建带有脸型风格适配器的风格扩散变换器,可自适应调整音色与韵律动态,同时与唇部感知语义特征保持同步。大量实验表明,ImaginTalk相比当前最优基线生成的语音在保真度、语义准确性及音色情感表达上均有显著提升。演示见项目页:https://imagintalk.github.io。
原文摘要 · Abstract (English)
Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting individuals with aphonia. Despite recent progress, existing methods struggle to achieve unified cross-modal alignment across semantics, timbre, and emotional prosody from visual cues, prompting us to propose Consistent Video-to-Speech (CV2S) as an extended task to enhance cross-modal consistency. To tackle emerging challenges, we introduce ImaginTalk, a novel cross-modal diffusion framework that generates faithful speech using only visual input, operating within a discrete space. Specifically, we propose a discrete lip aligner that predicts discrete speech tokens from lip videos to capture semantic information, while an error detector identifies misaligned tokens, which are subsequently refined through masked language modeling with BERT. To further enhance the expressiveness of the generated speech, we develop a style diffusion transformer equipped with a face-style adapter that adaptively customizes identity and prosody dynamics across both the channel and temporal dimensions while ensuring synchronization with lip-aware semantic features. Extensive experiments demonstrate that ImaginTalk can generate high-fidelity speech with more accurate semantic details and greater expressiveness in timbre and emotion compared to state-of-the-art baselines. Demos are shown at our project page: https://imagintalk.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。