arXiv:2506.01293cs.CVcs.AI2025-06被引 4

新基准M3STR测试模型理解视觉结构化知识能力

Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation

  • 用多模态知识图构建含结构关系的图像,测试模型抽象理解力
  • 26个顶尖模型在该任务上表现普遍不足,暴露推理短板
  • 适合关注多模态推理与知识融合的研究者使用

多模态大语言模型(MLLMs)将异构模态融入大语言模型,实现对多样化场景与对象的全面理解。尽管已有大量评测基准和排行榜,但多数忽略模型对以视觉形式呈现的结构化抽象世界知识的理解能力。为此,我们提出一种新的评估范式,并设计了M3STR基准,基于多模态地图的结构化理解。该基准利用多模态知识图生成包含子图架构与多模态实体的合成图像。M3STR要求模型不仅识别视觉输入中的多模态实体,还需解析其复杂的拓扑关系。我们描述了基准的统计特征与自动化构建流程,并对26个先进MLLM进行了广泛实证分析。结果揭示模型在处理抽象视觉结构化信息方面存在持续缺陷,为提升MLLM整体推理能力指明关键方向。代码与数据已开源。

原文摘要 · Abstract (English)

Multi-modal large language models (MLLMs) incorporate heterogeneous modalities into LLMs, enabling a comprehensive understanding of diverse scenarios and objects. Despite the proliferation of evaluation benchmarks and leaderboards for MLLMs, they predominantly overlook the critical capacity of MLLMs to comprehend world knowledge with structured abstractions that appear in visual form. To address this gap, we propose a novel evaluation paradigm and devise M3STR, an innovative benchmark grounded in the Multi-Modal Map for STRuctured understanding. This benchmark leverages multi-modal knowledge graphs to synthesize images encapsulating subgraph architectures enriched with multi-modal entities. M3STR necessitates that MLLMs not only recognize the multi-modal entities within the visual inputs but also decipher intricate relational topologies among them. We delineate the benchmark's statistical profiles and automated construction pipeline, accompanied by an extensive empirical analysis of 26 state-of-the-art MLLMs. Our findings reveal persistent deficiencies in processing abstractive visual information with structured knowledge, thereby charting a pivotal trajectory for advancing MLLMs' holistic reasoning capacities. Our code and data are released at https://github.com/zjukg/M3STR

多模态知识图谱评测基准推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。