统一模型实现个性化长时语音生成与多模态理解
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- 采用'脑-口'双轨架构,分离多模态推理与实时语音生成
- 支持流式零样本变声克隆,长时间保持音色稳定
- 数据效率高,适合需要个性化语音的智能应用
我们提出MGM-Omni,一种统一的全模态大语言模型,实现全模态理解与富有表现力的长时语音生成。不同于分步处理的流水线,MGM-Omni采用“脑-口”设计,通过双轨令牌化架构,清晰分离多模态推理与实时语音生成。该设计支持高效的跨模态交互和低延迟流式语音生成。在理解方面,统一训练策略结合双音频编码器设计,实现多样声学条件下长时音频感知。在生成方面,基于块的并行解码方案缩小了文本与语音令牌率差距,加速推理,并支持长时间稳定的零样本语音克隆。相比同期工作,MGM-Omni以更高效的数据训练实现上述能力。大量实验表明,MGM-Omni在保持长序列音色一致性、生成自然且情境相关的语音,以及提升长时音频与全模态理解方面,优于现有开源模型。MGM-Omni建立了一种高效端到端的全模态理解与可控个性化长时语音生成范式。
原文摘要 · Abstract (English)
We present MGM-Omni, a unified Omni LLM for omni-modal understanding and expressive, long-horizon speech generation. Unlike cascaded pipelines that isolate speech synthesis, MGM-Omni adopts a "brain-mouth" design with a dual-track, token-based architecture that cleanly decouples multimodal reasoning from real-time speech generation. This design enables efficient cross-modal interaction and low-latency, streaming speech generation. For understanding, a unified training strategy coupled with a dual audio encoder design enables long-form audio perception across diverse acoustic conditions. For generation, a chunk-based parallel decoding scheme narrows the text speech token-rate gap, accelerating inference and supporting streaming zero-shot voice cloning with stable timbre over extended durations. Compared to concurrent work, MGM-Omni achieves these capabilities with markedly data-efficient training. Extensive experiments demonstrate that MGM-Omni outperforms existing open source models in preserving timbre identity across extended sequences, producing natural and context-aware speech, and achieving superior long-form audio and omnimodal understanding. MGM-Omni establishes an efficient, end-to-end paradigm for omnimodal understanding and controllable, personalised long-horizon speech generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。