arXiv:2509.25229cs.AI2025-09被引 3

测试大模型从照片生成平面图的能力,发现多数表现还不如随机猜测。

Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models

  • 用50套公寓照片+每套20张内景图,评估模型还原空间布局能力
  • 多数模型性能低于随机水平,人类仍大幅领先,图像模型尤其不听指令
  • 首次量化对比不同模型的空间智能,适合关注AI认知能力的研究者

我们提出Blueprint-Bench,一个通过将公寓照片转换为准确2D平面图来评估AI模型空间推理能力的基准。尽管输入图片属于现代多模态模型的训练分布,但空间重建需要真正的空间智能:推断房间布局、理解连通性并保持一致尺度。我们在包含50套公寓、每套约20张室内图像的数据集上,评估了领先的语言模型(GPT-5、Claude 4 Opus、Gemini 2.5 Pro、Grok-4)、图像生成模型(GPT-Image、NanoBanana)和代理系统(Codex CLI、Claude Code)。评分算法基于生成与真实平面图的房间连接图和尺寸排序相似性。结果揭示当前AI能力存在显著盲区:多数模型表现处于或低于随机基线,而人类表现显著更优。图像生成模型在遵循指令方面尤其薄弱,而具有迭代优化能力的代理方法也未明显优于单次生成。Blueprint-Bench提供了首个跨模型架构比较空间智能的数值框架。我们将持续评估新模型,并欢迎社区提交,以监测通用AI系统中空间智能的出现。

原文摘要 · Abstract (English)

We introduce Blueprint-Bench, a benchmark designed to evaluate spatial reasoning capabilities in AI models through the task of converting apartment photographs into accurate 2D floor plans. While the input modality (photographs) is well within the training distribution of modern multimodal models, the task of spatial reconstruction requires genuine spatial intelligence: inferring room layouts, understanding connectivity, and maintaining consistent scale. We evaluate leading language models (GPT-5, Claude 4 Opus, Gemini 2.5 Pro, Grok-4), image generation models (GPT-Image, NanoBanana), and agent systems (Codex CLI, Claude Code) on a dataset of 50 apartments with approximately 20 interior images each. Our scoring algorithm measures similarity between generated and ground-truth floor plans based on room connectivity graphs and size rankings. Results reveal a significant blind spot in current AI capabilities: most models perform at or below a random baseline, while human performance remains substantially superior. Image generation models particularly struggle with instruction following, while agent-based approaches with iterative refinement capabilities show no meaningful improvement over single-pass generation. Blueprint-Bench provides the first numerical framework for comparing spatial intelligence across different model architectures. We will continue evaluating new models as they are released and welcome community submissions, monitoring for the emergence of spatial intelligence in generalist AI systems.

空间推理多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。