arXiv:2607.21072cs.CV2026-07被引 1

让图像生成模型直接在像素空间回答空间问题,更真实评估其空间认知能力。

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

论文配图:Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
图 1 · 摘自论文原文
  • 提出可适配任意基准的视觉评估框架,让图像模型以绘图等方式输出答案
  • 在470个样本的SpatialGen-Bench上验证,图像模型在直接像素表达任务中表现不俗
  • 适合研究视觉模型空间认知、对比文本与像素表达优劣的研究者

空间智能是智能体从静态语义理解迈向物理世界交互的关键。许多空间任务基于连续视觉场景,位置、区域和路径通过指向、标记或绘制比精确坐标或离散文本符号更自然表达。然而现有空间推理基准多要求坐标、选项或文本,造成图像生成模型的答案接口不匹配,难以与文本输出视觉语言模型在相同任务语义下评估。为此,我们提出ProVisE(协议化视觉评估)框架,能从图像生成模型获取受协议约束的视觉答案,并解析为兼容原始度量的结构化预测。ProVisE还包含一个代理构建器,可为新基准构建并验证特定任务协议。我们进一步推出SpatialGen-Bench,一个涵盖14个空间子任务、四个能力层级、多种答案形式的470样本诊断基准。在统一设置下评估代表性文本输出VLM与图像生成模型,并在六个外部空间基准上验证代理协议构建的有效性。结果表明:当空间答案可直接在像素空间外化时,图像生成模型表现具有竞争力;而在组合式空间推理上,文本输出VLM仍具明显优势。这些发现揭示了像素空间表达与文本推理的互补优势,建立了一个指标兼容的测试平台,用于研究图像生成模型的空间认知能力。

原文摘要 · Abstract (English)

Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.

空间认知图像生成评估框架视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。