arXiv:2501.01957cs.CVcs.SD2025-01NeurIPS被引 228

让AI同时听懂看懂,实现接近实时的视听交互。

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

  • 分阶段训练让大模型逐步掌握视觉与语音理解能力。
  • 无需独立语音识别和合成模块,响应速度大幅提升。
  • 适合需要低延迟视听交互的应用场景,如智能客服。

近期多模态大语言模型(MLLM)主要聚焦视觉与文本模态的融合,对语音在交互中的作用关注较少。然而语音在多模态对话系统中至关重要,而同时实现高性能视觉与语音任务仍因模态差异面临挑战。本文提出一种精心设计的多阶段训练方法,逐步引导大模型理解视觉与语音信息,最终实现流畅的视听交互。该方法不仅保留了强大的视觉-语言能力,还实现了无需独立语音识别(ASR)与语音合成(TTS)模块的端到端语音对话,显著提升多模态响应速度。在图像、视频与语音任务基准测试中,我们的模型展现出卓越的视觉与语音性能,支持近实时视听交互。代码已公开于 https://github.com/VITA-MLLM/VITA。

原文摘要 · Abstract (English)

Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in multimodal dialogue systems, and implementing high-performance in both vision and speech tasks remains a significant challenge due to the fundamental modality differences. In this paper, we propose a carefully designed multi-stage training methodology that progressively trains LLM to understand both visual and speech information, ultimately enabling fluent vision and speech interaction. Our approach not only preserves strong vision-language capacity, but also enables efficient speech-to-speech dialogue capabilities without separate ASR and TTS modules, significantly accelerating multimodal end-to-end response speed. By comparing our method against state-of-the-art counterparts across benchmarks for image, video, and speech tasks, we demonstrate that our model is equipped with both strong visual and speech capabilities, making near real-time vision and speech interaction. Code has been released at https://github.com/VITA-MLLM/VITA.

视听交互多模态实时响应大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。