首个支持多维度语音风格精细控制的大规模对话数据集,让语音交互更自然。
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
- 构建超830小时语音对话数据集,支持情绪、语速、音量等6维风格控制
- 微调后模型在风格控制任务中评分提升29.12%-42.33%,指令遵循率提高14.61%-40.09%
- 不仅提升语音表现力,还增强理解与推理能力,适合语音合成与对话系统研究者
当前语音对话模型缺乏细粒度语音风格控制能力,这严重影响了人机交互的自然性。为此,我们提出UltraVoice,首个专为多维度细粒度语音风格控制设计的大规模语音对话数据集。该数据集包含超过830小时的语音对话,覆盖情绪、语速、音量、口音、语言及复合风格共六个关键维度。在SLAM-Omni和VocalNet等领先模型上微调后,模型在风格控制任务上的表现显著提升:多维度控制任务中,平均意见分(MOS)提升29.12%-42.33%,指令遵循率(IFR)提升14.61%-40.09%。此外,在URO-Bench基准测试中,微调模型在基础与专业设置下分别实现+10.84%和+7.87%的平均性能提升。该数据集还可用于训练可调控文本转语音(TTS)模型,验证其高质量与广泛适用性。完整数据集与模型权重已开源。
原文摘要 · Abstract (English)
Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question answering. To address this limitation, we introduce UltraVoice, the first large-scale speech dialogue dataset engineered for multiple fine-grained speech style control. Encompassing over 830 hours of speech dialogues, UltraVoice provides instructions across six key speech stylistic dimensions: emotion, speed, volume, accent, language, and composite styles. Fine-tuning leading models such as SLAM-Omni and VocalNet on UltraVoice significantly enhances their fine-grained speech stylistic controllability without degrading core conversational abilities. Specifically, our fine-tuned models achieve improvements of 29.12-42.33% in Mean Opinion Score (MOS) and 14.61-40.09 percentage points in Instruction Following Rate (IFR) on multi-dimensional control tasks designed in the UltraVoice. Moreover, on the URO-Bench benchmark, our fine-tuned models demonstrate substantial gains in core understanding, reasoning, and conversational abilities, with average improvements of +10.84% on the Basic setting and +7.87% on the Pro setting. Furthermore, the dataset's utility extends to training controllable Text-to-Speech (TTS) models, underscoring its high quality and broad applicability for expressive speech synthesis. The complete dataset and model checkpoints are available at: https://github.com/bigai-nlco/UltraVoice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。