构建多模态评估基准,测试模型对音视频文本的协同理解能力
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
- 设计音视频强耦合任务,要求模型融合多模态信息
- 包含810段音视频同步数据,覆盖中英文共2617个问答对
- 引入细粒度视频定位任务,适合研究多模态理解的学者
本文提出OmniEval,一个面向多模态模型(如MiniCPM-O 2.6)的评估基准,涵盖视觉、听觉和文本输入。相比现有基准,OmniEval具有三大特点:(i) 全模态协同:设计强调音频与视频强耦合的任务,要求模型有效利用各模态的联合感知;(ii) 视频多样性:包含810段音视频同步视频,其中285段为中文,525段为英文;(iii) 任务多样且细粒度:共2617个问答对,含1412个开放式问题和1205个多选题,分为3大类12子类任务,新增名为Grounding的细粒度视频定位任务。我们在OmniEval上对多个多模态模型进行实验,旨在提供一个评估模型构建与理解多模态上下文连贯性的平台。代码与数据见https://omnieval-benchmark.github.io/。
原文摘要 · Abstract (English)
In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with existing benchmarks, our OmniEval has several distinctive features: (i) Full-modal collaboration: We design evaluation tasks that highlight the strong coupling between audio and video, requiring models to effectively leverage the collaborative perception of all modalities; (ii) Diversity of videos: OmniEval includes 810 audio-visual synchronized videos, 285 Chinese videos and 525 English videos; (iii) Diversity and granularity of tasks: OmniEval contains 2617 question-answer pairs, comprising 1412 open-ended questions and 1205 multiple-choice questions. These questions are divided into 3 major task types and 12 sub-task types to achieve comprehensive evaluation. Among them, we introduce a more granular video localization task named Grounding. Then we conduct experiments on OmniEval with several omni-modality models. We hope that our OmniEval can provide a platform for evaluating the ability to construct and understand coherence from the context of all modalities. Codes and data could be found at https://omnieval-benchmark.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。