arXiv:2503.01879cs.MMcs.CV2025-03被引 15

Nexus融合视听语言模态,实现高效多模态理解与生成。

Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision

  • 模块化架构支持灵活配置编码器-大模型-解码器组合。
  • 在视觉理解、语音问答等任务上超越基线模型,实测表现优异。
  • 适合需要跨模态交互的工业级多模态应用开发人员。

本文提出一个面向工业级应用的全模态大语言模型(LLM)流水线,整合音频、视觉与语言模态,以应对三模态数据稀缺、计算成本高及特征对齐复杂等挑战。该流水线包含三个核心组件:第一,模块化框架,支持多种编码器-大模型-解码器架构灵活配置;第二,轻量级训练策略,基于先进视觉-语言模型Qwen2.5-VL预训练音频-语言对齐,避免视觉模态昂贵的预训练;第三,音频合成流水线,从多样化真实场景生成高质量音文数据,支持自动语音识别与语音对话等应用。由此构建的工业级全模态模型Nexus经大量实验验证:(1)在视觉理解任务中,性能优于其基线模型Qwen2.5-VL-7B,证明训练策略高效;(2)在英文口语问答任务中,准确率超过同期对比模型MiniCPM-o2.6-7B(LLaMA Q基准);(3)在真实世界语音识别测试集上表现卓越,体现强鲁棒性;(4)在语音转文本翻译任务中,优于Qwen2-Audio-Instruct-7B;(5)在文本转语音任务中,基于预训练声码器(如Fishspeech1.4或CosyVoice2.0),在Seed-TTS基准上与基线声码器相当;(6)三模态对齐分析表明,引入音频模态显著增强了视觉与语言表征的一致性。

原文摘要 · Abstract (English)

This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limited tri-modal datasets, high computational costs, and complex feature alignments. Our pipeline consists of three main components: First, a modular framework enabling flexible configuration of various encoder-LLM-decoder architectures. Second, a lightweight training strategy that pre-trains audio-language alignment on the state-of-the-art vision-language model Qwen2.5-VL, thus avoiding the costly pre-training of vision-specific modalities. Third, an audio synthesis pipeline that generates high-quality audio-text data from diverse real-world scenarios, supporting applications such as Automatic Speech Recognition and Speech-to-Speech chat. To this end, we introduce an industry-level omni-modal LLM, Nexus. Extensive experiments validate the efficacy of our pipeline, yielding the following key findings:(1) In the visual understanding task, Nexus exhibits superior performance compared with its backbone model - Qwen2.5-VL-7B, validating the efficiency of our training strategy. (2) Within the English Spoken Question-Answering task, the model achieves better accuracy than the same-period competitor (i.e, MiniCPM-o2.6-7B) in the LLaMA Q. benchmark. (3) In our real-world ASR testset, Nexus achieves outstanding performance, indicating its robustness in real scenarios. (4) In the Speech-to-Text Translation task, our model outperforms Qwen2-Audio-Instruct-7B. (5) In the Text-to-Speech task, based on pretrained vocoder (e.g., Fishspeech1.4 or CosyVoice2.0), Nexus is comparable to its backbone vocoder on Seed-TTS benchmark. (6) An in-depth analysis of tri-modal alignment reveals that incorporating the audio modality enhances representational alignment between vision and language.

多模态语音生成大模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。