让语音合成适应环境,图文双条件驱动更真实声音生成
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
- 用文本+图像双条件控制语音生成,保持语义一致
- 在真实噪声环境下音质提升显著,多模态对齐更精准
- 适合虚拟人、智能助手等需环境感知的语音应用
我们提出 VoiceDiT,一种多模态生成模型,可从文本和视觉提示生成具有环境感知能力的语音与音频。尽管语音与文本对齐对可懂性至关重要,但在嘈杂环境下实现这一对齐仍是该领域中未被充分探索的重大挑战。为此,我们设计了一种名为 VoiceDiT 的新型音频生成流水线,包含三个核心组件:(1) 用于预训练的大规模合成语音数据集和用于微调的精细化真实世界语音数据集;(2) Dual-DiT 模型,能够高效保留语音对齐信息并准确反映环境条件;(3) 基于扩散模型的 Image-to-Audio Translator,使模型能桥接音频与图像,生成与多模态提示对齐的环境音。大量实验结果表明,VoiceDiT 在真实世界数据集上优于先前模型,在音频质量与模态融合方面均有显著提升。
原文摘要 · Abstract (English)
We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy conditions remains a significant and underexplored challenge in the field. To address this, we present a novel audio generation pipeline named VoiceDiT. This pipeline includes three key components: (1) the creation of a large-scale synthetic speech dataset for pre-training and a refined real-world speech dataset for fine-tuning, (2) the Dual-DiT, a model designed to efficiently preserve aligned speech information while accurately reflecting environmental conditions, and (3) a diffusion-based Image-to-Audio Translator that allows the model to bridge the gap between audio and image, facilitating the generation of environmental sound that aligns with the multi-modal prompts. Extensive experimental results demonstrate that VoiceDiT outperforms previous models on real-world datasets, showcasing significant improvements in both audio quality and modality integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。