让语音合成适应环境声音,生成更自然的音频场景。
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
- 用流匹配方法联合生成语音和背景音。
- 在无标注录音中自监督提取配对数据,解决训练数据不足。
- 可精细控制背景音量,适合虚拟助手、游戏音效等场景。
近期文本到语音(TTS)技术已实现高度自然的语音合成,但如何将语音与复杂背景环境融合仍具挑战。我们提出UmbraTTS,一种基于流匹配的TTS模型,能同时生成语音与环境音频,以文本和声学上下文为条件。该模型支持对背景音量的细粒度控制,生成多样、连贯且情境感知的音频场景。关键难点在于缺乏自然语境下语音与背景音对齐的数据。为此,我们设计了一种自监督框架,从无标注录音中提取语音、背景音频及对应文本。大量评估表明,UmbraTTS显著优于现有基线,生成自然、高质量且环境感知的音频。
原文摘要 · Abstract (English)
Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce UmbraTTS, a flow-matching based TTS model that jointly generates both speech and environmental audio, conditioned on text and acoustic context. Our model allows fine-grained control over background volume and produces diverse, coherent, and context-aware audio scenes. A key challenge is the lack of data with speech and background audio aligned in natural context. To overcome the lack of paired training data, we propose a self-supervised framework that extracts speech, background audio, and transcripts from unannotated recordings. Extensive evaluations demonstrate that UmbraTTS significantly outperformed existing baselines, producing natural, high-quality, environmentally aware audios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。