测试视觉语言模型能否识别并分类推理错误。
MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models
- 构建1997个含单一推理错误的多模态样本,覆盖24个子领域。
- 顶级模型仅66.65%正确识别错误类型,暴露模型理解短板。
- 适合关注模型可解释性与推理能力评估的研究者。
视觉语言模型(VLMs)在多模态学习中取得进展,引发其是否真正理解所处理内容的疑问。关键问题是:这些模型能否检测推理过程中的错误并识别错误类型?为此,我们提出MMErroR,一个包含1997个样本的多模态基准,每个样本嵌入一个连贯的推理错误。样本覆盖六个顶层领域下的24个子领域,确保广泛性和分类丰富性。不同于聚焦答案正确性的现有基准,MMErroR采用过程级、以错误为中心的评估方式,要求模型在视觉和语言上下文中检测错误并分类。我们评估了12个代表性VLMs,即使表现最佳的Gemini-3-Pro-Preview也仅在66.65%的情况下正确分类错误,凸显识别推理错误的挑战。准确识别错误的能力为多模态模型能力提供了重要洞察。
原文摘要 · Abstract (English)
Recent advances in Vision-Language Models (VLMs) have improved performance in multi-modal learning, raising the question of whether these models truly understand the content they process. Crucially, can VLMs detect when a reasoning process is wrong and identify its error type? To answer this, we present MMErroR, a multi-modal benchmark of 1997 samples, each embedding a single coherent reasoning error. These samples span 24 subdomains across six top-level domains, ensuring broad coverage and taxonomic richness. Unlike existing benchmarks that focus on answer correctness, MMErroR targets a process-level, error-centric evaluation that requires models to detect incorrect reasoning and classify the error type within both visual and linguistic contexts. We evaluate 12 representative VLMs, and even the best model, Gemini-3-Pro-Preview, classifies the error correctly in only 66.65\% of cases, underscoring the challenge of identifying erroneous reasoning. Furthermore, the ability to accurately identify errors offers valuable insights into the capabilities of multi-modal models. Project Page: https://mmerror-benchmark.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。