首个统一音频视频理解与生成的多模态大模型,支持同步指令驱动。
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
- 采用编码器-语言模型-解码器结构,融合时空音视频信息并引入同步感知查询
- 在20万条高质量指令数据上训练,复杂时序任务上超越现有模型
- 适合需要音视频协同理解与生成的研究者和开发者
本文提出JavisGPT,首个统一的多模态大语言模型(MLLM),用于联合音频-视频(JAV)理解与生成。该模型采用简洁的编码器-语言模型-解码器架构,包含用于时空音视频融合的SyncFusion模块和同步感知可学习查询,以连接预训练的JAV-DiT生成器。这一设计使模型能从多模态指令中实现时间一致的音视频理解与生成。我们设计了一个三阶段训练流程:多模态预训练、音视频微调及大规模指令微调,逐步构建多模态理解与生成能力。指令微调阶段构建了JavisInst-Omni数据集,包含超过20万条由GPT-4o标注的音视频文本对话,覆盖多样且多层次的理解与生成场景。在多个音视频理解与生成基准测试中,实验表明JavisGPT优于现有多模态大模型,尤其在复杂且需时间同步的任务中表现突出。
原文摘要 · Abstract (English)
This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation. JavisGPT has a concise encoder-LLM-decoder architecture, which has a SyncFusion module for spatio-temporal audio-video fusion and synchrony-aware learnable queries to bridge a pretrained JAV-DiT generator. This design enables temporally coherent video-audio understanding and generation from multimodal instructions. We design an effective three-stage training pipeline consisting of multimodal pretraining, audio-video fine-tuning, and large-scale instruction-tuning, to progressively build multimodal comprehension and generation from existing vision-language models. For instruction tuning, we construct JavisInst-Omni, a high-quality instruction dataset with over 200K GPT-4o-curated audio-video-text dialogues that cover diverse and multi-level comprehension and generation scenarios. On JAV comprehension and generation benchmarks, our experiments show that JavisGPT outperforms existing MLLMs, particularly in complex and temporally synchronized settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。