用多模型融合提升长视频音画理解能力,效果显著优于现有方法。
QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models
- 通过融合大模型、语言模型与语音识别,构建长视频音画理解新流程。
- 在VideoMME数据集上性能提升38.75%,超越VideoLlaMA2等先进模型。
- 适合需要分析长时视频内容的场景,如智能剪辑、视频摘要与机器人认知。
面向长视频音画理解任务,传统大视听模型多局限于几分钟内的短视频评估。本文提出QMAVIS(Q Team-多模态音视频智能理解系统),通过晚融合大型多模态模型(LMMs)、大语言模型(LLMs)与语音识别模型,构建全新的长视频音画理解流程。该系统填补了长视频分析的空白,尤其适用于数分钟至一小时以上的视频,拓展了智能理解、内容分析与具身智能等应用前景。在包含音频信息的长视频数据集VideoMME(含字幕)上,QMAVIS相比当前最佳模型VideoLlaMA2和InternVL2实现38.75%的性能提升;在PerceptionTest和EgoSchema等挑战性数据集上亦达到最高2%的改进,表现优异。定性实验表明,系统能精准捕捉长视频中不同场景的细节,并把握整体叙事脉络。消融实验证明各组件在融合管道中的关键作用。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) for video-audio understanding have traditionally been evaluated only on shorter videos of a few minutes long. In this paper, we introduce QMAVIS (Q Team-Multimodal Audio Video Intelligent Sensemaking), a novel long video-audio understanding pipeline built through a late fusion of LMMs, Large Language Models, and speech recognition models. QMAVIS addresses the gap in long-form video analytics, particularly for longer videos of a few minutes to beyond an hour long, opening up new potential applications in sensemaking, video content analysis, embodied AI, etc. Quantitative experiments using QMAVIS demonstrated a 38.75% improvement over state-of-the-art video-audio LMMs like VideoLlaMA2 and InternVL2 on the VideoMME (with subtitles) dataset, which comprises long videos with audio information. Evaluations on other challenging video understanding datasets like PerceptionTest and EgoSchema saw up to 2% improvement, indicating competitive performance. Qualitative experiments also showed that QMAVIS is able to extract the nuances of different scenes in a long video audio content while understanding the overarching narrative. Ablation studies were also conducted to ascertain the impact of each component in the fusion pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。