arXiv:2503.05223cs.SDcs.CV2025-03NAACL被引 4

无需音频提示,直接从视频生成语音并保留说话人特征。

DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility

  • 端到端模型直接从视频帧预测梅尔频谱图。
  • 在LRS2和LRS3数据集上同时提升语音可懂度与说话人特征保真度。
  • 适合语音合成、视频驱动语音生成方向的研究者参考。

视频到语音(V2S)合成任务旨在仅从静默视频输入中生成语音,因其需仅凭视觉线索精确重建语音内容与说话人特征,挑战性高于其他语音合成任务。近期,音视频预训练已使V2S不再依赖额外声学提示即可实现训练收敛。然而,现有方法仍难以在语音可懂度与说话人特征保留之间取得平衡。我们分析此局限并提出DiVISe(Direct Visual-Input Speech Synthesis),一个端到端的V2S模型,仅从视频帧直接预测梅尔频谱图。尽管不使用任何声学提示,DiVISe仍有效保留生成语音中的说话人特征,在LRS2和LRS3数据集上的客观与主观评估指标均表现更优。结果表明,DiVISe不仅在语音可懂度上优于现有模型,且在数据量与模型参数增加时具备更强扩展性。代码与权重见https://github.com/PussyCat0700/DiVISe。

原文摘要 · Abstract (English)

Video-to-speech (V2S) synthesis, the task of generating speech directly from silent video input, is inherently more challenging than other speech synthesis tasks due to the need to accurately reconstruct both speech content and speaker characteristics from visual cues alone. Recently, audio-visual pre-training has eliminated the need for additional acoustic hints in V2S, which previous methods often relied on to ensure training convergence. However, even with pre-training, existing methods continue to face challenges in achieving a balance between acoustic intelligibility and the preservation of speaker-specific characteristics. We analyzed this limitation and were motivated to introduce DiVISe (Direct Visual-Input Speech Synthesis), an end-to-end V2S model that predicts Mel-spectrograms directly from video frames alone. Despite not taking any acoustic hints, DiVISe effectively preserves speaker characteristics in the generated audio, and achieves superior performance on both objective and subjective metrics across the LRS2 and LRS3 datasets. Our results demonstrate that DiVISe not only outperforms existing V2S models in acoustic intelligibility but also scales more effectively with increased data and model parameters. Code and weights can be found at https://github.com/PussyCat0700/DiVISe.

视频生成语音合成端到端说话人保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。