arXiv:2502.04326cs.CVcs.AI2025-02中稿 · ICLR被引 140

首个评估多模态大模型真实世界理解能力的基准测试

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

  • 设计强耦合音视频任务,考验模型跨模态协同感知能力
  • 覆盖1662段同步音视频与3172个问答对,涵盖67个细分类别
  • 由80名专家多轮标注,适合研究真实场景理解的开发者使用

我们提出WorldSense,首个同时涵盖视觉、音频和文本输入的多模态视频理解评估基准。相较于现有基准,WorldSense具备三大特点:(i) 强耦合多模态设计,任务要求模型有效利用音视频间的协同感知;(ii) 视频与任务多样性,包含1,662段音视频同步视频,系统划分至8个主要领域和67个细粒度子类别,覆盖广泛真实场景,并设有3,172个多项选择题问答对,对应26种不同任务;(iii) 高质量标注,所有问答对均由80名专家经多轮校正人工标注,确保可靠性。基于此基准,我们对多种先进模型进行广泛评估,结果显示当前模型在真实场景理解上仍面临显著挑战(最佳准确率为65.1%)。通过分析模型局限性,旨在为提升真实世界理解能力提供指导。我们期望WorldSense能成为评估模型构建与理解多模态连贯语境能力的重要平台。

原文摘要 · Abstract (English)

We introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features: (i)collaboration of omni-modality, we design the evaluation tasks to feature a strong coupling of audio and video, requiring models to effectively utilize the synergistic perception of omni-modality; (ii)diversity of videos and tasks, WorldSense encompasses a diverse collection of 1,662 audio-visual synchronised videos, systematically categorized into 8 primary domains and 67 fine-grained subcategories to cover the broad scenarios, and 3,172 multi-choice QA pairs across 26 distinct tasks to enable the comprehensive evaluation; (iii)high-quality annotations, all the QA pairs are manually labeled by 80 expert annotators with multiple rounds of correction to ensure quality. Based on our WorldSense, we extensively evaluate various state-of-the-art models. The experimental results indicate that existing models face significant challenges in understanding real-world scenarios (65.1% best accuracy). By analyzing the limitations of current models, we aim to provide valuable insight to guide development of real-world understanding. We hope our WorldSense can provide a platform for evaluating the ability in constructing and understanding coherent contexts from omni-modality.

多模态视频理解评测基准真实世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。