用图像控制语音的沉浸式合成,让声音匹配场景空间感。
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
- 通过图像提示编码器将视觉场景融入语音生成流程。
- 结合混响分类与优化,使语音空间感与场景一致。
- 适合虚拟现实、游戏等需要沉浸体验的应用场景。
控制语音风格与特征对适应特定场景和用户需求至关重要。以往的文本到语音(TTS)研究主要关注如何生成自然流畅的语音,如语调、节奏和清晰度,但忽略了合成语音的空间感知,而空间感知在游戏和虚拟现实中能提供沉浸式体验。为此,本文提出一种新的多模态TTS方法——图像指示沉浸式语音合成(I2TTS)。具体而言,引入场景提示编码器,将视觉场景提示直接整合进合成流程以控制语音生成;同时提出混响分类与优化技术,调整合成的梅尔频谱图,增强沉浸感,确保混响条件准确匹配场景。实验结果表明,该模型在不牺牲语音自然度的前提下,实现了高质量的场景与空间匹配,标志着上下文感知语音合成领域的重要进展。
原文摘要 · Abstract (English)
Controlling the style and characteristics of speech synthesis is crucial for adapting the output to specific contexts and user requirements. Previous Text-to-speech (TTS) works have focused primarily on the technical aspects of producing natural-sounding speech, such as intonation, rhythm, and clarity. However, they overlook the fact that there is a growing emphasis on spatial perception of synthesized speech, which may provide immersive experience in gaming and virtual reality. To solve this issue, in this paper, we present a novel multi-modal TTS approach, namely Image-indicated Immersive Text-to-speech Synthesis (I2TTS). Specifically, we introduce a scene prompt encoder that integrates visual scene prompts directly into the synthesis pipeline to control the speech generation process. Additionally, we propose a reverberation classification and refinement technique that adjusts the synthesized mel-spectrogram to enhance the immersive experience, ensuring that the involved reverberation condition matches the scene accurately. Experimental results demonstrate that our model achieves high-quality scene and spatial matching without compromising speech naturalness, marking a significant advancement in the field of context-aware speech synthesis. Project demo page: https://spatialTTS.github.io/ Index Terms-Speech synthesis, scene prompt, spatial perception
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。