arXiv:2511.18448cs.CV2025-11被引 2

构建首个综合评估事件流多模态大模型的基准,覆盖理解、识别与空间推理。

EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs

  • 设计八类任务,支持从原始事件流直接推理的多模态模型评测
  • 包含超百万对事件-文本数据,支持大规模训练与评估
  • 首次引入3D空间推理任务,揭示当前模型在细粒度识别上的不足

多模态大语言模型(MLLM)在事件驱动视觉领域取得显著进展,但其能力在统一基准下的全面评估仍不充分。本文提出EventBench,一个包含八种多样化任务指标及大规模事件流数据集的基准。其四大特点:(1) 开放性,发布全部原始事件流与任务指令;(2) 任务多样性,涵盖理解、识别与空间推理任务;(3) 空间维度整合,首创3D空间推理任务;(4) 数据规模,配套训练集超一百万对事件-文本对。使用EventBench评估GPT-5、Gemini-2.5 Pro等闭源模型,以及Qwen2.5-VL、InternVL3等开源模型,还有直接处理原始事件流的EventGPT。结果表明,现有事件驱动型MLLM在事件流理解上表现良好,但在细粒度识别与空间推理方面仍存在明显短板。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have made significant advancements in event-based vision, yet the comprehensive evaluation of their capabilities within a unified benchmark remains largely unexplored. In this work, we introduce EventBench, a benchmark that offers eight diverse task metrics together with a large-scale event stream dataset. EventBench differs from existing event-based benchmarks in four key aspects: (1) openness in accessibility, releasing all raw event streams and task instructions across eight evaluation metrics; (2) diversity in task coverage, spanning understanding, recognition, and spatial reasoning tasks for comprehensive capability assessment; (3) integration in spatial dimensions, pioneering the design of 3D spatial reasoning tasks for event-based MLLMs; and (4) scale in data volume, with an accompanying training set of over one million event-text pairs supporting large-scale training and evaluation. Using EventBench, we evaluate state-of-the-art closed-source models such as GPT-5 and Gemini-2.5 Pro, leading open-source models including Qwen2.5-VL and InternVL3, and event-based MLLMs such as EventGPT that directly process raw event streams. Extensive evaluation reveals that while current event-based MLLMs demonstrate strong performance in event stream understanding, they continue to struggle with fine-grained recognition and spatial reasoning.

事件流多模态基准测试空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。