arXiv:2503.11026eess.AScs.CV2025-03ICCV

用条件流匹配保持说话人特征,实现跨语言音视频零样本转换

MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation

  • 采用双模态引导的条件流匹配,同步保留语音与面部特征
  • 引入x-vectors增强说话人一致性,提升零样本翻译效果
  • 适合多语言音视频生成、影视字幕对齐等场景

尽管文本转语音模型已有进展,音视频到音视频(AV2AV)转换仍面临关键挑战:保持原始与翻译后语音及面部特征的说话人一致性。为此,我们提出一种基于条件流匹配(CFM)的零样本音视频渲染器,利用音频与视觉模态的强双重引导。通过多模态引导的CFM,模型稳健地保留了说话人特有特征,并增强了零样本AV2AV转换能力。在音频方面,通过整合鲁棒的说话人嵌入(x-vectors)增强CFM过程,以强化说话人一致性;同时向面部生成模块传递情感细节。音频与视觉引导独立于语义或语言内容,使渲染器能有效处理不同语言中单语说话人的零样本转换任务。实验证明,基于面部信息条件化的高质量梅尔频谱图不仅提升了合成语音质量,还正向影响面部生成,整体在LSE和FID评分上取得改进。

原文摘要 · Abstract (English)

Despite recent advances in text-to-speech (TTS) models, audio-visual-to-audio-visual (AV2AV) translation still faces a critical challenge: maintaining speaker consistency between the original and translated vocal and facial features. To address this issue, we propose a conditional flow matching (CFM) zero-shot audio-visual renderer that utilizes strong dual guidance from both audio and visual modalities. By leveraging multimodal guidance with CFM, our model robustly preserves speaker-specific characteristics and enhances zero-shot AV2AV translation abilities. For the audio modality, we enhance the CFM process by integrating robust speaker embeddings with x-vectors, which serve to bolster speaker consistency. Additionally, we convey emotional nuances to the face rendering module. The guidance provided by both audio and visual cues remains independent of semantic or linguistic content, allowing our renderer to effectively handle zero-shot translation tasks for monolingual speakers in different languages. We empirically demonstrate that the inclusion of high-quality mel-spectrograms conditioned on facial information not only enhances the quality of the synthesized speech but also positively influences facial generation, leading to overall performance improvements in LSE and FID score. Our code is available at https://github.com/Peter-SungwooCho/MAVFlow.

音视频生成零样本说话人一致流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。