arXiv:2510.13344cs.SDcs.CL2025-10被引 6

统一语音与音乐生成,用动态专家网络解决数据不均衡问题。

UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE

  • 采用动态容量专家混合框架,按需分配计算资源。
  • 三阶段训练提升跨域协同,语音音乐生成均达顶尖水平。
  • 适合音频生成、多模态模型研究者参考。

当前统一多模态模型趋向于综合性内容生成,但音频领域仍面临挑战:语音与音乐常被孤立开发,导致通用音频合成进展受限。这一分离源于任务冲突与严重数据不平衡。为此,我们提出UniMoE-Audio,基于新型动态容量专家混合(Dynamic-Capacity MoE)框架的统一语音与音乐生成模型。其架构引入Top-P路由策略实现专家数量动态分配,并设计混合专家结构:路由专家捕捉领域特异性知识,共享专家处理通用特征,空专家实现自适应计算跳过。为应对数据不平衡,采用三阶段训练课程:1)独立专家训练,利用原始数据集在无干扰下建立各“原型专家”;2)MoE集成与预热,将专家嵌入模型,用平衡数据子集预热门控模块与共享专家;3)协同联合训练,在完整平衡数据集上端到端优化,增强跨域协同。大量实验表明,UniMoE-Audio不仅在主流语音与音乐生成基准上达到领先性能,还显著缓解了传统联合训练中的性能下降问题。结果证明,专业化专家架构与精心设计的训练策略对推动通用音频生成具有巨大潜力。

原文摘要 · Abstract (English)

Recent advances in unified multimodal models indicate a clear trend towards comprehensive content generation. However, the auditory domain remains a significant challenge, with music and speech often developed in isolation, hindering progress towards universal audio synthesis. This separation stems from inherent task conflicts and severe data imbalances, which impede the development of a truly unified audio generation model. To address this challenge, we propose UniMoE-Audio, a unified speech and music generation model within a novel Dynamic-Capacity Mixture-of-Experts (MoE) framework. Architecturally, UniMoE-Audio introduces a Top-P routing strategy for dynamic expert number allocation, and a hybrid expert design comprising routed experts for domain-specific knowledge, shared experts for domain-agnostic features, and null experts for adaptive computation skipping. To tackle data imbalance, we introduce a three-stage training curriculum: 1) Independent Specialist Training leverages original datasets to instill domain-specific knowledge into each "proto-expert" without interference; 2) MoE Integration and Warmup incorporates these specialists into the UniMoE-Audio architecture, warming up the gate module and shared expert using a subset of balanced dataset; and 3) Synergistic Joint Training trains the entire model end-to-end on the fully balanced dataset, fostering enhanced cross-domain synergy. Extensive experiments show that UniMoE-Audio not only achieves state-of-the-art performance on major speech and music generation benchmarks, but also demonstrates superior synergistic learning, mitigating the performance degradation typically seen in naive joint training. Our findings highlight the substantial potential of specialized MoE architecture and curated training strategies in advancing the field of universal audio generation. Homepage: https://mukioxun.github.io/Uni-MoE-site/home.html

音频生成MoE语音合成音乐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。