arXiv:2505.14404cs.CV2025-05被引 10

构建自由式视觉中间态评估基准,更真实测试多模态大模型的渐进推理能力。

ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations

  • 设计自由风格视觉中间态生成流程,支持动态适应任务需求。
  • 在4类任务上验证18个模型,发现现有模型在复杂推理中表现有限。
  • 提供可扩展评估框架,适合研究视觉链式思考机制的学者使用。

视觉交错思维(VI-CoT)使多模态大语言模型能基于逐步的视觉中间状态(IVS)持续更新理解与决策空间,类似人类思考过程,在多个任务中展现出显著成效。然而,当前基准普遍采用固定形式的IVS,可能扭曲原始思维轨迹,无法真实评估模型的内在推理能力。更重要的是,现有研究未系统探究IVS对非受限推理性能的影响因素。为此,我们提出专用于评估的基准ViC-Bench,包含迷宫导航、拼图、具身长周期规划及复杂计数四类代表性任务,每项均配备支持自适应函数调用的自由风格IVS生成管道。为系统评估VI-CoT能力,我们设计三阶段渐进式评估策略与新指标体系,并引入增量提示信息注入策略以分析提示因素影响。我们对18个先进多模态大模型进行了全面评测,揭示了其在该能力上的关键洞察。ViC-Bench已公开于Huggingface。

原文摘要 · Abstract (English)

Visual-Interleaved Chain-of-Thought (VI-CoT) enables Multi-modal Large Language Models (MLLMs) to continually update their understanding and decision space based on step-wise intermediate visual states (IVS), much like a human would, which has demonstrated impressive success in various tasks, thereby leading to emerged advancements in related downstream benchmarks. Despite promising progress, current benchmarks provide models with relatively fixed IVS, rather than free-style IVS, whch might forcibly distort the original thinking trajectories, failing to evaluate their intrinsic reasoning capabilities. More importantly, existing benchmarks neglect to systematically explore the impact factors that IVS would impart to the untamed reasoning performance. To tackle above gaps, we introduce a specialized benchmark termed ViC-Bench, consisting of four representive tasks, i.e., maze navigation, jigsaw puzzle, embodied long-horizon planning, as well as complex counting, where each task has dedicated free-style IVS generation pipeline supporting adaptive function calls. To systematically examine VI-CoT capability, we propose a thorough evaluation suite incorporating a progressive three-stage strategy with targeted new metrics. Besides, we establish Incremental Prompting Information Injection strategy to ablatively explore the prompting factors for VI-CoT. We extensively conduct evaluations for 18 advanced MLLMs, revealing key insights into their VI-CoT capability. The introduced ViC-Bench has been made publicly available at Huggingface.

多模态链式思考推理评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。