FleSpeech支持多模态提示,灵活控制语音风格与音色
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
- 通过多模态提示编码器统一处理文本、音频和视觉提示
- 可同时保留指定说话人音色并调整语音风格
- 适合需要角色化语音生成的创意场景
可控语音生成方法通常依赖单一或固定提示,限制了创作灵活性。在特定场景下,如需在保持选定说话人音色的同时调整语音风格,或根据角色视觉形象生成匹配语音时,现有方法难以满足需求。为此,我们提出FleSpeech,一种新型多阶段语音生成框架,通过整合多种形式的控制信号实现更灵活的语音属性操控。FleSpeech采用多模态提示编码器,将不同类型的文本、音频和视觉提示统一为一致表示,提升语音合成的适应性,并支持对生成语音的精准创造性控制。此外,我们构建了多模态数据集采集流程,以推动该领域研究与应用发展。主观与客观实验均验证了FleSpeech的有效性。音频样例可访问 https://kkksuper.github.io/FleSpeech/
原文摘要 · Abstract (English)
Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjusting the style while preserving a selected speaker's timbre, or choosing a style and generating a voice that matches a character's visual appearance. To overcome these challenges, we propose \textit{FleSpeech}, a novel multi-stage speech generation framework that allows for more flexible manipulation of speech attributes by integrating various forms of control. FleSpeech employs a multimodal prompt encoder that processes and unifies different text, audio, and visual prompts into a cohesive representation. This approach enhances the adaptability of speech synthesis and supports creative and precise control over the generated speech. Additionally, we develop a data collection pipeline for multimodal datasets to facilitate further research and applications in this field. Comprehensive subjective and objective experiments demonstrate the effectiveness of FleSpeech. Audio samples are available at https://kkksuper.github.io/FleSpeech/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。