arXiv:2501.13707cs.CVcs.AI2025-01被引 14

首个面向事件流的多模态大模型,提升语义理解能力

EventVL: Understand Event Streams via Multimodal Large Language Model

  • 构建140万组事件-图像-文本数据集,打通多模态语义鸿沟
  • 设计时空表示与动态对齐机制,显著提升事件流语义表达
  • 在事件描述生成任务上超越现有基线,适合事件视觉研究者

事件驱动的视觉-语言模型在实际视觉任务中取得进展,但多数仅使用CLIP处理传统感知任务,难以充分理解事件流中的语义与上下文。为此,我们提出EventVL,首个面向事件流的生成式多模态大语言模型框架,实现显式语义理解。首先,我们构建包含约140万组高质量事件-图像/视频-文本配对的数据集,覆盖驾驶场景、人体运动等多样化场景,支持跨模态有效学习。其次,设计事件时空表示,通过多样化聚合与分割充分挖掘事件流信息。为进一步压缩语义空间,引入动态语义对齐机制,优化事件稀疏语义空间。大量实验表明,EventVL在事件描述生成和场景描述任务中显著优于现有多模态大模型基线。本研究有望推动事件视觉领域的发展。

原文摘要 · Abstract (English)

The event-based Vision-Language Model (VLM) recently has made good progress for practical vision tasks. However, most of these works just utilize CLIP for focusing on traditional perception tasks, which obstruct model understanding explicitly the sufficient semantics and context from event streams. To address the deficiency, we propose EventVL, the first generative event-based MLLM (Multimodal Large Language Model) framework for explicit semantic understanding. Specifically, to bridge the data gap for connecting different modalities semantics, we first annotate a large event-image/video-text dataset, containing almost 1.4 million high-quality pairs of data, which enables effective learning across various scenes, e.g., drive scene or human motion. After that, we design Event Spatiotemporal Representation to fully explore the comprehensive information by diversely aggregating and segmenting the event stream. To further promote a compact semantic space, Dynamic Semantic Alignment is introduced to improve and complete sparse semantic spaces of events. Extensive experiments show that our EventVL can significantly surpass existing MLLM baselines in event captioning and scene description generation tasks. We hope our research could contribute to the development of the event vision community.

事件视觉多模态大模型语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。