arXiv:2602.02537cs.CVcs.LG2026-02被引 4

测试大模型对视觉常识的记憶能力,区分记忆与推理。

WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

  • 分离视觉知识检索与推理,专注评估模型记忆
  • 覆盖从常见到罕见物体的分层分类体系
  • 适合评估模型事实性与幻觉率,适配前沿模型

我们提出WorldVQA,一个用于评估多模态大语言模型(MLLM)原子级视觉世界知识的基准。不同于现有评估常将视觉知识检索与推理混为一谈,WorldVQA 将二者解耦,严格测量模型“记住了什么”。该基准在分层分类体系中评估模型对视觉实体的定位与命名能力,涵盖从常见类别到长尾稀有物的广泛范围。我们期望WorldVQA能成为衡量视觉事实性的严格测试,为当前及下一代前沿模型的百科广度与幻觉率提供标准评估依据。

原文摘要 · Abstract (English)

We introduce WorldVQA, a benchmark designed to evaluate the atomic visual world knowledge of Multimodal Large Language Models (MLLMs). Unlike current evaluations, which often conflate visual knowledge retrieval with reasoning, WorldVQA decouples these capabilities to strictly measure "what the model memorizes." The benchmark assesses the atomic capability of grounding and naming visual entities across a stratified taxonomy, spanning from common head-class objects to long-tail rarities. We expect WorldVQA to serve as a rigorous test for visual factuality, thereby establishing a standard for assessing the encyclopedic breadth and hallucination rates of current and next-generation frontier models.

视觉知识大模型评测事实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。