arXiv:2501.04561cs.CLcs.CV2025-01被引 13

开源多模态大模型实现高效对齐与实时情感语音合成

OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis

  • 分两阶段训练:先对齐视觉与语音,再优化语音生成
  • 7B模型在多项评测超越更大模型,性能提升4点
  • 支持毫秒级实时语音输出,情感识别准确率提高7.7%

近期多模态学习进展虽提升了图像、文本和语音的理解与生成能力,但主要局限于专有模型。高质量多模态数据集匮乏及实时情感语音合成难题严重制约了开源研究。为此,我们提出 ame,一种两阶段训练框架,融合多模态对齐与语音生成,打造先进开源多模态大模型。在对齐阶段,预训练语音模型在图文任务上继续训练,实现从视觉到语音的(近)零样本泛化,优于基于三模态数据集训练的模型。在语音生成阶段,采用轻量解码器并结合直接偏好优化,实现高保真实时情感语音合成。实验表明, ame 在多模态、视觉-语言及语音-语言基准测试中均超越现有领先模型。其在 OmniBench 上较领先开源模型 VITA 提升 4 个百分点,仅用 5 倍少的训练样本与更小模型规模(7B vs. 7x8B)。此外, ame 在非自回归模式下实现 <1 秒延迟的实时语音生成,推理时间较自回归方法降低 5 倍,情感分类准确率提升 7.7%。

原文摘要 · Abstract (English)

Recent advancements in omnimodal learning have significantly improved understanding and generation across images, text, and speech, yet these developments remain predominantly confined to proprietary models. The lack of high-quality omnimodal datasets and the challenges of real-time emotional speech synthesis have notably hindered progress in open-source research. To address these limitations, we introduce \name, a two-stage training framework that integrates omnimodal alignment and speech generation to develop a state-of-the-art omnimodal large language model. In the alignment phase, a pre-trained speech model undergoes further training on text-image tasks, enabling (near) zero-shot generalization from vision to speech, outperforming models trained on tri-modal datasets. In the speech generation phase, a lightweight decoder is trained on speech tasks with direct preference optimization, enabling real-time emotional speech synthesis with high fidelity. Experiments show that \name surpasses state-of-the-art models across omnimodal, vision-language, and speech-language benchmarks. It achieves a 4-point absolute improvement on OmniBench over the leading open-source model VITA, despite using 5x fewer training samples and a smaller model size (7B vs. 7x8B). Additionally, \name achieves real-time speech generation with <1s latency at non-autoregressive mode, reducing inference time by 5x compared to autoregressive methods, and improves emotion classification accuracy by 7.7\%

多模态语音合成开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。