arXiv:2606.06538cs.CV2026-06被引 2

构建视觉多样性强的多模态推理评测集,揭示现有模型在真实场景下的理解短板。

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

  • 基于跨领域视觉概念分类体系,从搜索与数据集筛选海量图像。
  • 15个主流多模态大模型平均仅达64.0%准确率,最强模型仍不足七成。
  • 专为挑战模型视觉泛化能力设计,适合评估真实世界应用性能。

在真实应用中,模型需在多样场景下稳定表现。然而,现有多数多模态基准虽扩展任务类型,却未涵盖开放视觉输入所需的视觉多样性。本文提出WorldBench,一个具有挑战性且视觉多样性高的多模态推理评测基准,用于评估多模态大语言模型(MLLMs)。我们构建了跨多个领域(如生物体)数千个视觉概念的分类体系,据此从搜索引擎和现有数据集中筛选广泛图像,全面覆盖视觉世界。通过结构化试错,人工设计出前沿MLLM无法解答的难题。定量与人工评估显示,WorldBench的视觉多样性超越所有现有基准。对15个MLLM的评测揭示其视觉理解缺陷:最强模型仅达64.0%准确率,部分模型表现略高于随机水平。我们希望本工作凸显视觉多样性在构建多模态基准中的重要性。

原文摘要 · Abstract (English)

In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual diversity needed to handle open-ended visual inputs. We present WorldBench, a challenging and visually diverse reasoning benchmark to evaluate Multimodal Large Language Models (MLLMs). We build a taxonomy of thousands of visual concepts across multiple domains (e.g., living things). Guided by this taxonomy, we curate a broad collection of images from search engines and existing datasets to comprehensively represent the visual world. Through structured trial-and-error, we manually design challenging questions that frontier MLLMs fail to answer. On quantitative and human evaluations, WorldBench achieves higher visual diversity than any existing diverse benchmark. Evaluating 15 MLLMs on WorldBench reveals weaknesses in visual understanding: even the strongest model reaches only 64.0% accuracy, while some models perform marginally above chance-level. We hope our work highlights the importance of visual diversity in building multimodal benchmarks.

多模态评测基准视觉理解模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。