arXiv:2608.07584cs.CV2026-08

测试视觉语言模型在需验证决策任务中的表现,发现多数模型准确率不足40%。

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

论文配图:ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making
图 1 · 摘自论文原文
  • 构建39个领域场景下的390个可验证决策任务
  • 仅GPT-5.6-Sol达到75.6%的决策通过率,其余均低于40%
  • 视觉呈现相同但结构化输入显著提升性能,说明决策瓶颈存在

视觉语言模型(VLMs)在图像感知方面进展迅速,正支持越来越多依赖图像的现实任务。然而,许多任务不仅要求识别图像内容,更需利用视觉证据做出满足全局约束的完整决策。我们提出COMPLEXITYWORLD,一个包含39个领域启发式视觉世界和29类决策的390个任务的基准。每个任务由隐藏的结构化规范生成,渲染为视觉场景,并由可执行验证器评分,接受任何可行解。直接推理下,除GPT-5.6-Sol外所有模型的验证通过率(VAR)均低于40%,而后者达75.6%。当决策信息以结构化形式显式提供时,性能显著提升,但不同等效视觉呈现间差异明显。代理框架带来较小、模型依赖性增益。结果揭示了仅靠额外推理无法消除的视觉到决策瓶颈。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such tasks, however, require more than rec- ognizing what an image contains: a model must use visual evidence to make a complete decision whose parts jointly satisfy global constraints. We introduce COMPLEXITYWORLD, a benchmark of 390 tasks across 39 domain-inspired visual worlds and 29 decision categories. Each task is generated from a hidden structured specification, rendered as a visual scene, and scored by an exe- cutable verifier that accepts any feasible solution. Under direct inference, all evaluated models ex- cept GPT-5.6-Sol remain below 40% verifier ac- ceptance rate (VAR), while GPT-5.6-Sol reaches 75.6%. Performance improves substantially when the same decision information is made explicit in structured form, yet varies sharply across equiva- lent visual presentations. Agent scaffolds provide smaller, model-dependent gains. Together, these results reveal a persistent visual-to-decision bot- tleneck that additional inference alone does not remove.

视觉语言模型决策评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。