用自然语言统一控制语音与音乐生成,支持多语种表达式输出。
InstructAudio: Unified speech and music generation with natural language instruction
- 采用标准化指令-音素输入格式,结合联合与单扩散变压器层。
- 在50K小时语音与20K小时音乐数据上训练,实现跨模态对齐。
- 首个支持自然语言指令的语音音乐统一生成框架,适合多场景应用。
文本到语音(TTS)和文本到音乐(TTM)模型在基于指令的控制方面存在显著局限。TTS系统通常依赖参考音频来确定音色,仅提供有限的文本级属性控制,且很少支持对话生成;而TTM系统则受限于需专家知识标注的输入条件。这两种任务输入控制条件高度异质,难以联合建模。尽管两者共享共同的声学建模特性,但长期独立发展,尚未实现通过自然语言指令的统一建模。本文提出InstructAudio,一个统一框架,可基于自然语言描述(指令)控制音色(性别、年龄)、副语言特征(情绪、风格、口音)及音乐属性(流派、乐器、节奏、氛围)。该框架支持中英文的富有表现力语音、音乐及对话生成。模型采用联合与单扩散变压器层,使用标准化指令-音素输入格式,在50,000小时语音与20,000小时音乐数据上进行训练,实现多任务学习与跨模态对齐。图1展示了与主流TTS和TTM模型的性能对比,表明InstructAudio在多数指标上达到最优。据我们所知,InstructAudio是首个实现指令控制的语音与音乐统一生成框架。音频样例可访问:https://qiangchunyu.github.io/InstructAudio/
原文摘要 · Abstract (English)
Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only limited text-level attribute control, and rarely support dialogue generation. TTM systems are constrained by input conditioning requirements that depend on expert knowledge annotations. The high heterogeneity of these input control conditions makes them difficult to joint modeling with speech synthesis. Despite sharing common acoustic modeling characteristics, these two tasks have long been developed independently, leaving open the challenge of achieving unified modeling through natural language instructions. We introduce InstructAudio, a unified framework that enables instruction-based (natural language descriptions) control of acoustic attributes including timbre (gender, age), paralinguistic (emotion, style, accent), and musical (genre, instrument, rhythm, atmosphere). It supports expressive speech, music, and dialogue generation in English and Chinese. The model employs joint and single diffusion transformer layers with a standardized instruction-phoneme input format, trained on 50K hours of speech and 20K hours of music data, enabling multi-task learning and cross-modal alignment. Fig. 1 visualizes performance comparisons with mainstream TTS and TTM models, demonstrating that InstructAudio achieves optimal results on most metrics. To our best knowledge, InstructAudio represents the first instruction-controlled framework unifying speech and music generation. Audio samples are available at: https://qiangchunyu.github.io/InstructAudio/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。