提升音视频对齐精度,让多模态大模型听懂视频内容
Aligned Better, Listen Better for Audio-Visual Large Language Models
- 通过时空同步对齐音视频特征,实现细粒度融合
- 在520万数据上训练,显著降低幻觉并提升理解能力
- 适合音视频理解、指令生成等多模态任务研究者
音频是多模态视频理解的关键。视频本身包含音频,为视觉提供互补信息,且视频大语言模型常面临以音频为中心的任务场景。然而现有视频大模型与音视频大模型在利用音频信息方面存在不足,导致理解能力弱和幻觉问题。为此,我们从模型架构与数据集两方面入手:(1) 提出细粒度音视频大模型Dolphin,通过时序与空间维度的同步对齐,实现全面准确的视频理解。具体地,设计音视频多尺度适配器实现空间对齐;提出音视频交错融合策略实现时序对齐。(2) 构建音视频图文与指令微调数据集AVU,包含520万条多样化的开放性数据元组(视频、音频、问题、答案),并引入新型数据划分策略。大量实验表明,本模型不仅在音视频理解任务中表现优异,还有效缓解了潜在幻觉。
原文摘要 · Abstract (English)
Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric settings. However, existing Video-LLMs and Audio-Visual Large Language Models (AV-LLMs) exhibit deficiencies in exploiting audio information, leading to weak understanding and hallucinations. To solve the issues, we delve into the model architecture and dataset. (1) From the architectural perspective, we propose a fine-grained AV-LLM, namely Dolphin. The concurrent alignment of audio and visual modalities in both temporal and spatial dimensions ensures a comprehensive and accurate understanding of videos. Specifically, we devise an audio-visual multi-scale adapter for multi-scale information aggregation, which achieves spatial alignment. For temporal alignment, we propose audio-visual interleaved merging. (2) From the dataset perspective, we curate an audio-visual caption and instruction-tuning dataset, called AVU. It comprises 5.2 million diverse, open-ended data tuples (video, audio, question, answer) and introduces a novel data partitioning strategy. Extensive experiments show our model not only achieves remarkable performance in audio-visual understanding, but also mitigates potential hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。