Nexus融合视听语言模态,实现高效多模态理解与生成。
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
- 模块化架构支持灵活配置编码器-大模型-解码器组合。
- 在视觉理解、语音问答等任务上超越基线模型,实测表现优异。
- 适合需要跨模态交互的工业级多模态应用开发人员。
本文提出一个面向工业级应用的全模态大语言模型(LLM)流水线,整合音频、视觉与语言模态,以应对三模态数据稀缺、计算成本高及特征对齐复杂等挑战。该流水线包含三个核心组件:第一,模块化框架,支持多种编码器-大模型-解码器架构灵活配置;第二,轻量级训练策略,基于先进视觉-语言模型Qwen2.5-VL预训练音频-语言对齐,避免视觉模态昂贵的预训练;第三,音频合成流水线,从多样化真实场景生成高质量音文数据,支持自动语音识别与语音对话等应用。由此构建的工业级全模态模型Nexus经大量实验验证:(1)在视觉理解任务中,性能优于其基线模型Qwen2.5-VL-7B,证明训练策略高效;(2)在英文口语问答任务中,准确率超过同期对比模型MiniCPM-o2.6-7B(LLaMA Q基准);(3)在真实世界语音识别测试集上表现卓越,体现强鲁棒性;(4)在语音转文本翻译任务中,优于Qwen2-Audio-Instruct-7B;(5)在文本转语音任务中,基于预训练声码器(如Fishspeech1.4或CosyVoice2.0),在Seed-TTS基准上与基线声码器相当;(6)三模态对齐分析表明,引入音频模态显著增强了视觉与语言表征的一致性。
原文摘要 · Abstract (English)
This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limited tri-modal datasets, high computational costs, and complex feature alignments. Our pipeline consists of three main components: First, a modular framework enabling flexible configuration of various encoder-LLM-decoder architectures. Second, a lightweight training strategy that pre-trains audio-language alignment on the state-of-the-art vision-language model Qwen2.5-VL, thus avoiding the costly pre-training of vision-specific modalities. Third, an audio synthesis pipeline that generates high-quality audio-text data from diverse real-world scenarios, supporting applications such as Automatic Speech Recognition and Speech-to-Speech chat. To this end, we introduce an industry-level omni-modal LLM, Nexus. Extensive experiments validate the efficacy of our pipeline, yielding the following key findings:(1) In the visual understanding task, Nexus exhibits superior performance compared with its backbone model - Qwen2.5-VL-7B, validating the efficiency of our training strategy. (2) Within the English Spoken Question-Answering task, the model achieves better accuracy than the same-period competitor (i.e, MiniCPM-o2.6-7B) in the LLaMA Q. benchmark. (3) In our real-world ASR testset, Nexus achieves outstanding performance, indicating its robustness in real scenarios. (4) In the Speech-to-Text Translation task, our model outperforms Qwen2-Audio-Instruct-7B. (5) In the Text-to-Speech task, based on pretrained vocoder (e.g., Fishspeech1.4 or CosyVoice2.0), Nexus is comparable to its backbone vocoder on Seed-TTS benchmark. (6) An in-depth analysis of tri-modal alignment reveals that incorporating the audio modality enhances representational alignment between vision and language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。