用文字生成对口型人脸,无需音频也能精准同步
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
- 以发音单位(viseme)为桥梁,将文本转为可控的嘴型序列
- 逐步替换真实音频为合成音频,支持有无音频场景
- 结合关键点引导渲染,生成高保真对口型视频
生成语义一致且视觉准确的说话人脸,需弥合语言意义与面部动作之间的差距。尽管音频驱动方法仍占主导,但其依赖高质量音视频配对数据,且声学到唇动映射存在固有模糊性,限制了可扩展性和鲁棒性。为此,我们提出Text2Lip,一个以发音单位(viseme)为核心的框架,通过将文本输入嵌入结构化的发音序列,构建可解释的语音-视觉桥梁。这些中层单元作为唇动预测的语言基础先验。此外,我们设计了一种基于课程学习的渐进式发音-音频替换策略,使模型能逐步从真实音频过渡到由增强发音特征通过跨模态注意力重建的伪音频,从而在有音频和无音频场景下均实现鲁棒生成。最后,采用关键点引导的渲染器合成高保真面部视频,确保唇动精准同步。大量评估表明,Text2Lip在语义保真度、视觉真实性和模态鲁棒性方面优于现有方法,确立了可控且灵活的说话人脸生成新范式。
原文摘要 · Abstract (English)
Generating semantically coherent and visually accurate talking faces requires bridging the gap between linguistic meaning and facial articulation. Although audio-driven methods remain prevalent, their reliance on high-quality paired audio visual data and the inherent ambiguity in mapping acoustics to lip motion pose significant challenges in terms of scalability and robustness. To address these issues, we propose Text2Lip, a viseme-centric framework that constructs an interpretable phonetic-visual bridge by embedding textual input into structured viseme sequences. These mid-level units serve as a linguistically grounded prior for lip motion prediction. Furthermore, we design a progressive viseme-audio replacement strategy based on curriculum learning, enabling the model to gradually transition from real audio to pseudo-audio reconstructed from enhanced viseme features via cross-modal attention. This allows for robust generation in both audio-present and audio-free scenarios. Finally, a landmark-guided renderer synthesizes photorealistic facial videos with accurate lip synchronization. Extensive evaluations show that Text2Lip outperforms existing approaches in semantic fidelity, visual realism, and modality robustness, establishing a new paradigm for controllable and flexible talking face generation. Our project homepage is https://plyon1.github.io/Text2Lip/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。