arXiv:2506.16020cs.SDeess.AS2025-06中稿 · Interspeech 2025

用图像空间信息生成与场景匹配的立体声歌声。

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge

  • 通过视觉特征增强语音文本编码,融合空间信息。
  • 单步生成立体声歌声,且与场景视角一致。
  • 首个统一处理立体声歌声与音视同步的框架。

为探索图像中空间线索在生成带混响立体声歌声方面的优势,我们提出VS-Singer,一个基于视觉引导的模型,可从场景图像生成带有房间混响的立体声歌声。该模型包含三个模块:首先,模态交互网络将空间特征融入文本编码,生成富含空间信息的语言表征;其次,解码器采用一致性Schrödinger桥实现单步采样生成;此外,我们引入SFE模块提升音视匹配的一致性。据我们所知,这是首个在统一框架中结合立体声歌声合成与视觉声学匹配的研究。实验表明,VS-Singer可在单步内有效生成与场景视角一致的立体声歌声。

原文摘要 · Abstract (English)

To explore the potential advantages of utilizing spatial cues from images for generating stereo singing voices with room reverberation, we introduce VS-Singer, a vision-guided model designed to produce stereo singing voices with room reverberation from scene images. VS-Singer comprises three modules: firstly, a modal interaction network integrates spatial features into text encoding to create a linguistic representation enriched with spatial information. Secondly, the decoder employs a consistency Schrödinger bridge to facilitate one-step sample generation. Moreover, we utilize the SFE module to improve the consistency of audio-visual matching. To our knowledge, this study is the first to combine stereo singing voice synthesis with visual acoustic matching within a unified framework. Experimental results demonstrate that VS-Singer can effectively generate stereo singing voices that align with the scene perspective in a single step.

语音合成视觉引导立体声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。