arXiv:2510.23594cs.CV2025-10被引 5

评测大模型视觉推理能力,能发现思考过程中的错误。

PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection

  • 设计谜题类视觉任务,要求识别思维链中首个错误步骤。
  • 顶尖模型虽能生成流畅回答,却常无法发现简单逻辑错误。
  • 适合研究可信多模态模型的开发者和评估者使用。

多模态大语言模型在视觉-语言任务上取得显著进展,但其推理过程仍不可靠。我们提出PRISM-Bench,一个基于谜题的视觉任务基准,不仅评估模型能否解题,更关注其推理过程。不同于以往仅衡量最终答案准确率,PRISM-Bench引入诊断任务:给定一个视觉谜题和包含恰好一处错误的分步思维链(CoT),模型需找出第一个错误步骤。该设置可实现对逻辑一致性、错误检测与视觉推理能力的细粒度评估。谜题涉及多步符号、几何与类比推理,避免依赖表层模式匹配的捷径。对主流多模态大模型的测试显示,模型在生成流畅思维链时,常无法识别简单的逻辑错误。通过将答案生成与推理验证分离,PRISM-Bench为多模态推理能力提供了更清晰的评估视角,凸显了构建诊断性评估协议对发展可信多模态模型的重要性。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have achieved remarkable progress on vision-language tasks, yet their reasoning processes remain sometimes unreliable. We introduce PRISM-Bench, a benchmark of puzzle-based visual challenges designed to evaluate not only whether models can solve problems, but how their reasoning unfolds. Unlike prior evaluations that measure only final-answer accuracy, PRISM-Bench introduces a diagnostic task: given a visual puzzle and a step-by-step chain-of-thought (CoT) containing exactly one error, models must identify the first incorrect step. This setting enables fine-grained assessment of logical consistency, error detection, and visual reasoning. The puzzles in PRISM-Bench require multi-step symbolic, geometric, and analogical reasoning, resisting shortcuts based on superficial pattern matching. Evaluations across state-of-the-art MLLMs reveal a persistent gap between fluent generation and faithful reasoning: models that produce plausible CoTs often fail to locate simple logical faults. By disentangling answer generation from reasoning verification, PRISM-Bench offers a sharper lens on multimodal reasoning competence and underscores the need for diagnostic evaluation protocols in the development of trustworthy MLLMs.

多模态推理评测错误检测视觉谜题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。