Lyra让AI更懂语音,实现高效多模态认知。
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
- 用开源模型+多模态LoRA降低训练成本
- 1.5百万多模态数据+12千长语音样本提升语音理解
- 兼顾性能与效率,适合语音交互场景
随着多模态大语言模型(MLLMs)的发展,突破单一领域能力限制以满足更通用、高效的AI需求变得至关重要。然而,以往的全模态模型对语音的探索不足,未能充分整合语音与多模态信息。我们提出Lyra,一种高效且以语音为中心的MLLM,增强了包括长语音理解、声音感知、跨模态效率和无缝语音交互在内的多模态能力。为实现高效性与语音中心特性,Lyra采用三项策略:(1) 利用现有开源大模型及提出的多模态LoRA,降低训练成本与数据需求;(2) 使用潜在多模态正则化器与提取器,强化语音与其他模态之间的关联,从而提升模型表现;(3) 构建高质量、大规模数据集,包含150万组多模态(语言、视觉、音频)数据样本和1.2万条长语音样本,使Lyra能够处理复杂长语音输入并实现更鲁棒的全认知能力。相比其他全模态方法,Lyra在多种视觉-语言、视觉-语音和语音-语言基准上达到顶尖性能,同时使用更少的计算资源与训练数据。
原文摘要 · Abstract (English)
As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond single-domain capabilities is essential to meet the demands for more versatile and efficient AI. However, previous omni-models have insufficiently explored speech, neglecting its integration with multi-modality. We introduce Lyra, an efficient MLLM that enhances multimodal abilities, including advanced long-speech comprehension, sound understanding, cross-modality efficiency, and seamless speech interaction. To achieve efficiency and speech-centric capabilities, Lyra employs three strategies: (1) leveraging existing open-source large models and a proposed multi-modality LoRA to reduce training costs and data requirements; (2) using a latent multi-modality regularizer and extractor to strengthen the relationship between speech and other modalities, thereby enhancing model performance; and (3) constructing a high-quality, extensive dataset that includes 1.5M multi-modal (language, vision, audio) data samples and 12K long speech samples, enabling Lyra to handle complex long speech inputs and achieve more robust omni-cognition. Compared to other omni-methods, Lyra achieves state-of-the-art performance on various vision-language, vision-speech, and speech-language benchmarks, while also using fewer computational resources and less training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。