Tutti实现可调控的多歌手合唱生成,支持动态声线与音色纹理建模。
Tutti: Expressive Multi-Singer Synthesis via Structure-Level Timbre Control and Vocal Texture Modeling
- 通过结构感知歌手提示实现多歌手灵活调度
- 引入互补纹理学习,提升合唱声学真实感
- 适合音乐创作、虚拟合唱等复杂多声部场景
现有歌唱语音合成系统虽能生成高保真独唱,但在全局音色控制下难以处理歌曲中动态的多歌手编排及声部纹理。为此,我们提出Tutti统一框架,引入结构感知歌手提示,使歌手调度随音乐结构动态变化;并提出条件引导的变分自编码器进行互补纹理学习,捕捉空间混响、频谱融合等隐式声学纹理。实验表明,Tutti在精确多歌手调度和显著提升合唱声学真实感方面表现优异,为复杂多歌手编排提供了新范式。音频样例见https://annoauth123-ctrl.github.io/Tutii_Demo/
原文摘要 · Abstract (English)
While existing Singing Voice Synthesis systems achieve high-fidelity solo performances, they are constrained by global timbre control, failing to address dynamic multi-singer arrangement and vocal texture within a single song. To address this, we propose Tutti, a unified framework designed for structured multi-singer generation. Specifically, we introduce a Structure-Aware Singer Prompt to enable flexible singer scheduling evolving with musical structure, and propose Complementary Texture Learning via Condition-Guided VAE to capture implicit acoustic textures (e.g., spatial reverberation and spectral fusion) that are complementary to explicit controls. Experiments demonstrate that Tutti excels in precise multi-singer scheduling and significantly enhances the acoustic realism of choral generation, offering a novel paradigm for complex multi-singer arrangement. Audio samples are available at https://annoauth123-ctrl.github.io/Tutii_Demo/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。