arXiv:2511.05516cs.CLcs.AI2025-11被引 22

首个统一语音理解生成与编辑的模型,支持自然语言指令自由修改语音。

Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation

  • 采用连续语音分词器MingTok-Audio,融合语义与声学特征统一表示。
  • 在12项指标中8项达新SOTA,中文语音克隆WER低至0.95。
  • 首个支持自然语言指令的通用语音编辑模型,无需时间戳约束。

现有语音模型在理解与生成任务中对词元表示的需求相互冲突,阻碍了基于指令的自由形式语音编辑。为此,我们提出一种新框架,统一语音理解、生成与编辑。核心是首个有效融合语义与声学特征的连续语音分词器MingTok-Audio,适用于两类任务。基于此,我们构建了语音语言模型Ming-UniAudio,实现了生成与理解能力的平衡,在ContextASR基准上8项指标达到新SOTA。尤其在中文语音克隆中,实现0.95的Seed-TTS-WER。进一步训练出专用语音编辑模型Ming-UniAudio-Edit,首次实现仅依赖自然语言指令的通用自由形式语音编辑,可同时处理语义与声学修改,无需时间戳条件。为严谨评估编辑能力,我们引入Ming-Freeform-Audio-Edit,首个专用于指令式自由形式语音编辑的基准,涵盖多样场景与多维评估维度(语义正确性、声学质量、指令对齐)。我们开源了连续语音分词器、统一基础模型及指令式编辑模型,推动统一语音理解、生成与操控的发展。

原文摘要 · Abstract (English)

Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech language models from performing instruction-based free-form editing. To solve this challenge, we introduce a novel framework that unifies speech understanding, generation, and editing. The core of our unified model is a unified continuous speech tokenizer MingTok-Audio, the first continuous tokenizer to effectively integrate semantic and acoustic features, which makes it suitable for both understanding and generation tasks. Based on this unified continuous audio tokenizer, we developed the speech language model Ming-UniAudio, which achieved a balance between generation and understanding capabilities. Ming-UniAudio sets new state-of-the-art (SOTA) records on 8 out of 12 metrics on the ContextASR benchmark. Notably, for Chinese voice cloning, it achieves a highly competitive Seed-TTS-WER of 0.95. Leveraging this foundational model, we further trained a dedicated speech editing model Ming-UniAudio-Edit, the first speech language model that enables universal, free-form speech editing guided solely by natural language instructions, handling both semantic and acoustic modifications without timestamp condition. To rigorously assess the editing capability and establish a foundation for future research, we introduce Ming-Freeform-Audio-Edit, the first comprehensive benchmark tailored for instruction-based free-form speech editing, featuring diverse scenarios and evaluation dimensions spanning semantic correctness, acoustic quality, and instruction alignment. We open-sourced the continuous audio tokenizer, the unified foundational model, and the free-form instruction-based editing model to facilitate the development of unified audio understanding, generation, and manipulation.

语音生成语音编辑大模型统一建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。