让语音合成同时控制音色和环境声,实现真实场景下的零样本语音生成。
Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling

- 分离语音、噪声与混响三要素,用独立提示分别控制音色与环境。
- 在模拟数据上训练,实测生成语音自然度高、环境还原准。
- 适合需要灵活调整语音环境的应用,如虚拟助手、语音克隆。
现有零样本语音合成系统虽能生成自然且保真度高的语音,但通常依赖高质量说话人提示,并将环境信息与说话人特征混淆,限制了实际应用。本文提出扩展版DAIEN-TTS,一种环境感知的零样本语音合成框架,通过解耦并联合建模语音、背景噪声与混响,实现通过独立的说话人提示与环境提示对音色与声学环境进行分别控制。基于流匹配的F5-TTS构建,引入语音-环境分离模块,将带环境的语音分解为语音、噪声与混响成分,并注入扩散变换器以实现环境感知生成。训练使用混合干净语音与噪声及房间冲激响应的模拟数据,结合跨说话人条件策略抑制环境分支中的说话人信息泄露。当有真实数据时,可进一步微调以弥合模拟到真实的域差距。推理阶段采用三重无分类器引导机制实现对语音、噪声与混响的细粒度控制,并通过信噪比自适应策略使合成语音与环境提示对齐。在模拟与真实测试集上的实验表明,DAIEN-TTS能生成具有高自然度、强说话人相似性以及精确噪声与混响还原的个性化语音,且可控性超越现有环境感知语音合成系统。
原文摘要 · Abstract (English)
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or entangle the acoustic environment with speaker characteristics, limiting their real-world applicability. We present an extended DAIEN-TTS, an environment-aware zero-shot TTS framework that disentangles and jointly models speech, background noise, and reverberation, enabling independent control over timbre and acoustic environment through separate speaker and environment prompts. Built upon the flow-matching-based F5-TTS, it uses a speech-environment separation module to decompose environmental speech into speech, noise, and reverberation components, which are injected into the Diffusion Transformer for environment-aware generation. Training uses simulated data constructed by mixing clean speech with noise and room impulse responses, together with a cross-speaker conditioning strategy that suppresses speaker information leakage from the environment branch. When real-world data are available, the system can be further fine-tuned to bridge the simulated-to-real domain gap.At inference, a triple classifier-free guidance mechanism enables fine-grained control over speech, noise, and reverberation, and a signal-to-noise-ratio adaptation strategy aligns the synthesized speech with the environment prompt. Experiments on simulated and real-world test sets show that DAIEN-TTS generates environmental personalized speech with high naturalness, strong speaker similarity, and faithful noise and reverberation reproduction, while offering controllability beyond prior environment-aware TTS systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。