首个开源语音编辑模型,支持情绪风格等多维度精准控制。
Step-Audio-EditX Technical Report
- 仅用大规模合成数据训练,无需额外模块或嵌入先验
- 在情感编辑等任务中优于MiniMax-2.6-hd和Doubao-Seed-TTS-2.0
- 适合需要高表达性语音编辑的研究与应用开发者
我们提出Step-Audio-EditX,首个基于大语言模型的开源音频模型,可在情感、语调风格及副语言特征上实现富有表现力且可迭代的音频编辑,并具备强大的零样本文本到语音(TTS)能力。其核心创新在于仅使用大规模合成数据进行训练,避免了对嵌入式先验或辅助模块的依赖。这种大间隔学习方法使模型同时具备迭代控制能力和丰富的语音表现力,标志着从传统表示层解耦关注点的根本转变。评估结果表明,Step-Audio-EditX在情感编辑及其他细粒度控制任务中超越MiniMax-2.6-hd和Doubao-Seed-TTS-2.0。
原文摘要 · Abstract (English)
We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguistics alongside robust zero-shot text-to-speech (TTS) capabilities. Our core innovation lies in leveraging only large-margin synthetic data, which circumvents the need for embedding-based priors or auxiliary modules. This large-margin learning approach enables both iterative control and high expressivity across voices, and represents a fundamental pivot from the conventional focus on representation-level disentanglement. Evaluation results demonstrate that Step-Audio-EditX surpasses both MiniMax-2.6-hd and Doubao-Seed-TTS-2.0 in emotion editing and other fine-grained control tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。