首个面向长视频的多模态事件理解基准,支持时序精细分析。
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
- 自动构建多模态视频筛选与事件边界检测流程
- 涵盖105K个事件,覆盖8.4K段高质量长视频
- 适合研究多模态视频理解与视频LLM的学者使用
尽管视频理解取得了显著进展,但多数工作仍局限于粗粒度或仅视觉的任务。现实世界视频包含视觉、音频和语音等多模态信息,并由一系列事件构成连贯的叙事。缺乏带细粒度事件标注的多模态视频数据,以及人工标注成本高,是实现全面多模态视频感知的主要障碍。为此,我们提出一个自动流程:包括高质量多模态视频筛选、语义一致的多模态事件边界检测,以及跨模态相关性感知的事件描述生成。基于此,我们构建了首个视觉-音频-语言-事件理解基准LongVALE,包含105,000个具有精确时间边界的多模态事件,分布在8,400段高质量长视频中,并提供关系感知的详细描述。此外,我们建立基线模型,首次使视频大语言模型(LLMs)能够进行多模态细粒度时间视频理解。大量实验验证了LongVALE在推动全面多模态视频理解方面的有效性与巨大潜力。
原文摘要 · Abstract (English)
Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks. However, real-world videos encompass omni-modal information (vision, audio, and speech) with a series of events forming a cohesive storyline. The lack of multi-modal video data with fine-grained event annotations and the high cost of manual labeling are major obstacles to comprehensive omni-modality video perception. To address this gap, we propose an automatic pipeline consisting of high-quality multi-modal video filtering, semantically coherent omni-modal event boundary detection, and cross-modal correlation-aware event captioning. In this way, we present LongVALE, the first-ever Vision-Audio-Language Event understanding benchmark comprising 105K omni-modal events with precise temporal boundaries and detailed relation-aware captions within 8.4K high-quality long videos. Further, we build a baseline that leverages LongVALE to enable video large language models (LLMs) for omni-modality fine-grained temporal video understanding for the first time. Extensive experiments demonstrate the effectiveness and great potential of LongVALE in advancing comprehensive multi-modal video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。