评测大模型对视频中物体变化的动态理解能力,发现现有模型短板并提供新基准。
$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models
- 构建双视频对比任务,评估模型对同一场景前后状态变化的理解。
- 包含270个场景、2932道题,覆盖50+子任务,测试4项核心能力。
- 适合研究视觉推理、动态感知和多模态理解的学者使用。
现代大型多模态模型在静态图像和单状态时空理解方面表现出色,但对同一空间场景中两个不同视频观测间物体动态变化的推理能力仍缺乏探索。这种在一致环境中推理变化的能力对空间智能发展至关重要。本文提出M^3-Verse,一个面向多模态、多状态、多维度的基准测试,基于成对视频提供室内场景变化前后的多视角观测。该基准包含270个场景和2932个问题,涵盖50多个子任务,用于检验4项核心能力。我们评估了16个先进LMMs,发现其在追踪状态转换方面存在局限。为此,我们提出一种简单有效的基线方法,在多状态感知上取得显著提升。M^3-Verse为下一代具备更全面动态视觉理解能力的模型研发提供了挑战性新平台。数据与构建流程见https://github.com/Wal-K-aWay/M3-Verse_pipeline和https://www.modelscope.cn/datasets/WalKaWay/M3-Verse。
原文摘要 · Abstract (English)
Modern Large Multimodal Models (LMMs) have demonstrated extraordinary ability in static image and single-state spatial-temporal understanding. However, their capacity to comprehend the dynamic changes of objects within a shared spatial context between two distinct video observations, remains largely unexplored. This ability to reason about transformations within a consistent environment is particularly crucial for advancements in the field of spatial intelligence. In this paper, we introduce $M^3-Verse$, a Multi-Modal, Multi-State, Multi-Dimensional benchmark, to formally evaluate this capability. It is built upon paired videos that provide multi-perspective observations of an indoor scene before and after a state change. The benchmark contains a total of 270 scenes and 2,932 questions, which are categorized into over 50 subtasks that probe 4 core capabilities. We evaluate 16 state-of-the-art LMMs and observe their limitations in tracking state transitions. To address these challenges, we further propose a simple yet effective baseline that achieves significant performance improvements in multi-state perception. $M^3-Verse$ thus provides a challenging new testbed to catalyze the development of next-generation models with a more holistic understanding of our dynamic visual world. You can get the construction pipeline from https://github.com/Wal-K-aWay/M3-Verse_pipeline and full benchmark data from https://www.modelscope.cn/datasets/WalKaWay/M3-Verse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。