arXiv:2608.13613eess.AScs.LG2026-08

用统一模型实现多样语音生成与编辑,支持真实与虚构声音设计。

VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation

论文配图:VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation
图 1 · 摘自论文原文
  • 构建混合数据管道,融合信号处理与生成模型扩充语音多样性。
  • 提出改进的扩散变换器,提升复杂条件下的多任务性能。
  • 可灵活编辑语音情感、音调等属性,适合语音创意与个性化应用。

近期生成模型的突破使文本到语音(TTV)成为可能,能够直接根据文本描述合成语音。然而,现有系统面临两大挑战:一是难以生成涵盖真实人类发音者和虚构角色的多样化声音;二是缺乏鲁棒且灵活的语音编辑能力,如语音克隆及情感、语调等属性修改。本文提出 VoiceDesigner,一个统一的文本到语音生成与编辑框架,支持多样且可控的声音设计。为解决上述问题,我们从两个方面提出方案:首先,开发一种混合数据管道,结合数字信号处理技术与语音生成模型,构建覆盖真实与虚构声音的多样化语音数据集;其次,引入架构优化的扩散变压器,更好处理复杂条件并增强多任务表现,实现统一的语音生成与编辑。通过主观与客观评估,VoiceDesigner 在语音描述与编辑指令对齐方面表现更优,同时在感知质量与语音可用性上媲美当前最先进模型。

原文摘要 · Abstract (English)

Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions. However, existing systems face two key challenges. First, they struggle to generate a diverse range of voices, spanning real-world human speakers and fictional characters. Second, they lack robust and flexible voice editing capabilities, such as voice cloning and the ability to modify attributes like emotion and tone. In this paper, we propose VoiceDesigner, a unified framework for text-to-voice generation and editing that supports diverse and controllable voice design. To tackle the above challenges, we propose solutions from two perspectives. First, we develop a hybrid data pipeline that leverages digital signal processing techniques and speech generation models to construct a diverse voice dataset covering both real-world and fictional voices. Second, we introduce a diffusion transformer with architectural improvements to better handle complex conditioning and enhance multi-task performance, enabling unified voice generation and editing. Through subjective and objective evaluations, VoiceDesigner achieves superior prompt alignment with both voice descriptions and editing instructions, while maintaining competitive perceptual quality and voice usability compared to state-of-the-art TTV models.

语音生成扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。