arXiv:2512.11995cs.CVcs.AI2025-12被引 1

评测视觉模型如何一步步探索并推理复杂问题。

V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions

  • 用多步提问链模拟视觉推理过程,拆解探索与回答能力。
  • 发现主流模型在规划和执行步骤上表现差异大,整体仍不足。
  • 适合研究多步视觉推理、评估模型思维过程的学者使用。

尽管许多视觉语言模型(VLMs)能回答明确、直接的问题,但在实际中面对需要多轮探索与推理的开放任务时表现不佳。这类任务要求在视觉空间中逐步探索与验证,类似人工智能侦探的思考路径,能生成更优答案。然而,中间步骤的庞大搜索空间使评估困难。为此,我们构建了「视觉多步探索推理基准」(V-REX),包含需原生多步探索的挑战性任务与评估协议。V-REX将多步探索推理转化为提问链(CoQ),并分离评估模型的(1)规划能力:选择探索性问题链;(2)执行能力:按顺序回答问题以收集信息得出最终答案。通过每步限制有限选项,实现对中间步骤的可靠定量与细粒度分析。评估SOTA专有及开源VLMs后,揭示出一致的规模增长趋势,规划与执行能力显著差异,以及多步探索推理仍有巨大提升空间。

原文摘要 · Abstract (English)

While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle in practice with complex open-ended tasks, which usually require multiple rounds of exploration and reasoning in the visual space. Such visual thinking paths not only provide step-by-step exploration and verification as an AI detective but also produce better interpretations of the final answers. However, these paths are challenging to evaluate due to the large exploration space of intermediate steps. To bridge the gap, we develop an evaluation suite, ``Visual Reasoning with multi-step EXploration (V-REX)'', which is composed of a benchmark of challenging visual reasoning tasks requiring native multi-step exploration and an evaluation protocol. V-REX covers rich application scenarios across diverse domains. V-REX casts the multi-step exploratory reasoning into a Chain-of-Questions (CoQ) and disentangles VLMs' capability to (1) Planning: breaking down an open-ended task by selecting a chain of exploratory questions; and (2) Following: answering curated CoQ sequentially to collect information for deriving the final answer. By curating finite options of questions and answers per step, V-REX achieves a reliable quantitative and fine-grained analysis of the intermediate steps. By assessing SOTA proprietary and open-sourced VLMs, we reveal consistent scaling trends, significant differences between planning and following abilities, and substantial room for improvement in multi-step exploratory reasoning.

视觉推理多步探索评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。