arXiv:2507.07644cs.AI2025-07被引 15

测试大模型对室内布局的空间推理能力,发现其常忽略物理约束。

FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations

  • 用JSON/XML结构化表示房间布局,设计空间任务评测
  • 模型在简单查询上表现尚可,但常违反空间逻辑
  • 适合研究空间推理或具身智能的学者参考

我们提出FloorplanQA,一个基于结构化室内场景表示(如厨房、客厅、卧室、浴室等)的大语言模型空间推理诊断基准。该基准采用符号化编码的JSON或XML布局,涵盖距离测量、可视性判断、路径规划和受限空间内的物体摆放等核心空间任务。对多种前沿开源与商用大模型的评估显示,尽管模型在浅层查询中表现良好,但在处理物理约束和保持空间一致性方面存在明显缺陷,尽管对微小空间扰动仍具鲁棒性。结果揭示了当前大模型在室内布局推理中的盲区:空间逻辑不一致。我们希望该基准能推动具备准确空间与几何推断能力的语言模型研究。

原文摘要 · Abstract (English)

We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large language models (LLMs). FloorplanQA is grounded in structured representations of indoor scenes, such as (e.g., kitchens, living rooms, bedrooms, bathrooms, and others), encoded symbolically in JSON or XML layouts. The benchmark covers core spatial tasks, including distance measurement, visibility, path finding, and object placement within constrained spaces. Our results across a variety of frontier open-source and commercial LLMs reveal that while models may succeed in shallow queries, they often fail to respect physical constraints, preserve spatial coherence, though they remain mostly robust to small spatial perturbations. FloorplanQA uncovers a blind spot in today's LLMs: inconsistent reasoning about indoor layouts. We hope this benchmark inspires new work on language models that can accurately infer and manipulate spatial and geometric properties in practical settings.

空间推理大模型评测结构化表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。