arXiv:2608.20414cs.AIcs.CV2026-08

提出新基准StateSight,测试大模型从图像重建空间结构的能力。

StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

论文配图:StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models
图 1 · 摘自论文原文
  • 用程序生成图像任务,分离空间推理与语言理解
  • 大模型在三类任务中准确率最高仅59.3%,远低于人类80.8%
  • 发现模型常犯空间状态重建错误,即使输出格式正确

视觉语言模型在多模态问答中应用日益广泛,但其从单张图像中重构潜在空间结构的能力难以独立评估。现有基准常混合感知、文字识别、领域知识、语言先验和推理能力。我们提出StateSight,一个程序生成的基准,涵盖立方体对面面推理、遮挡立方体塔计数和4-邻域连通组件计数三类任务。每类任务包含300个单图像提示,具有确定性真值标签和精确匹配评分。OpenAI GPT-5.5(API标识gpt-5.5)在三类任务上准确率分别为59.3%、33.3%和28.3%;Claude Sonnet 5为53.3%、18.7%和7.3%。所有最终运行均无格式错误。30名参与者在60个样本上的人类基线在各项任务中均优于模型,平均准确率为80.8%、68.8%和64.3%。可见推导分析揭示了图像状态重构和推理过程中的重复错误。我们还引入StateSight-Steps,包含900个交错图文样本和3600个确定性中间视觉状态。结果表明,格式正确的响应可能掩盖对可验证视觉推理所需空间结构的恢复失败。

原文摘要 · Abstract (English)

Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct latent spatial structure from a single image remains difficult to isolate. Broad benchmarks often combine perception, optical character recognition, domain knowledge, linguistic priors, and reasoning in the same evaluation. We introduce StateSight, a procedurally generated benchmark for cube-net opposite-face reasoning, occluded cube-tower counting, and 4-neighbor connected-component counting. Each task family contains 300 single-image prompts with deterministic oracle labels and exact-match scoring. OpenAI GPT-5.5, using the API model identifier gpt-5.5, achieved 59.3%, 33.3%, and 28.3% accuracy across the three tasks, while Claude Sonnet 5 achieved 53.3%, 18.7%, and 7.3%. All final direct runs had zero format errors. A 30-participant human baseline on 60 items exceeded both models on every task, with mean accuracies of 80.8%, 68.8%, and 64.3%. Visible-derivation analysis identified recurring errors in image-state reconstruction and reasoning procedure. We also introduce StateSight-Steps, a companion dataset of 900 interleaved image-text examples and 3,600 deterministic intermediate visual states. The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.

多模态空间推理模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。