首个长视频音画融合数据集,提升模型对复杂视听信息的理解能力。
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
- 构建58,000条音画指令数据集,支持长视频理解任务。
- 提出SAVEnVideo模型,在零样本任务上性能领先3.61%。
- 设计AVBench评测基准,挑战模型在长视频中的跨模态推理能力。
为提升长视频理解能力,现有视觉语言模型(Video-LLMs)仍面临有效整合丰富多样的音画信息的挑战。为此,我们提出:(i) 首个包含超过58,000条音画指令的长视频音画数据集SAVEn-Vid;(ii) 一种时间感知的音画大语言模型SAVEnVideo,基于该数据集进行微调;(iii) 设计了包含2,500个问答的AVBench评测基准,用于评估模型在长视频中对复杂音画交互的理解能力。实验表明,当前音画模型存在局限性;而SAVEnVideo在零样本长视频任务(Video-MME)上优于最佳视频模型3.61%,在零样本音画任务(Music-AVQA)上优于领先音画模型1.29%。在7B参数规模下,达到当前最优性能。相关数据与代码将在论文接收后公开。
原文摘要 · Abstract (English)
Endeavors have been made to explore Large Language Models for video analysis (Video-LLMs), particularly in understanding and interpreting long videos. However, existing Video-LLMs still face challenges in effectively integrating the rich and diverse audio-visual information inherent in long videos, which is crucial for comprehensive understanding. This raises the question: how can we leverage embedded audio-visual information to enhance long video understanding? Therefore, (i) we introduce SAVEn-Vid, the first-ever long audio-visual video dataset comprising over 58k audio-visual instructions. (ii) From the model perspective, we propose a time-aware Audio-Visual Large Language Model (AV-LLM), SAVEnVideo, fine-tuned on SAVEn-Vid. (iii) Besides, we present AVBench, a benchmark containing 2,500 QAs designed to evaluate models on enhanced audio-visual comprehension tasks within long video, challenging their ability to handle intricate audio-visual interactions. Experiments on AVBench reveal the limitations of current AV-LLMs. Experiments also demonstrate that SAVEnVideo outperforms the best Video-LLM by 3.61% on the zero-shot long video task (Video-MME) and surpasses the leading audio-visual LLM by 1.29% on the zero-shot audio-visual task (Music-AVQA). Consequently, at the 7B parameter scale, SAVEnVideo can achieve state-of-the-art performance. Our dataset and code will be released at https://ljungang.github.io/SAVEn-Vid/ upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。