构建三维空间认知评估基准,测试模型在复杂环境中的空间推理与规划能力。
Spatial Competence Benchmark

- 设计三级递进任务,要求模型输出可被确定性验证器或模拟器检验的执行结果。
- 前沿模型在任务阶梯上准确率逐级下降,低算力预算下性能提升迅速且很快饱和。
- 发现错误多源于局部几何合理但全局约束破坏,适合研究空间推理与规划的学者使用。
空间能力指维持环境一致内部表征,并据此推断离散结构、在约束条件下规划行为的能力。现有大模型的空间评估仅限于通过3D变换或视觉问答探测孤立原型。我们提出空间能力基准(SCBench),涵盖三个层级的能力任务,其输出需经确定性检查器或基于模拟器的评估器验证。在SCBench上,三个前沿模型的准确率随能力层级提升而单调下降。通过设定输出词元上限的实验表明,性能提升集中在低预算阶段且迅速饱和,失败主要源于局部看似合理但违反全局约束的几何构造。我们开源了任务生成器、验证器及可视化工具。
原文摘要 · Abstract (English)
Spatial competence is the quality of maintaining a consistent internal representation of an environment and using it to infer discrete structure and plan actions under constraints. Prevailing spatial evaluations for large models are limited to probing isolated primitives through 3D transformations or visual question answering. We introduce the Spatial Competence Benchmark (SCBench), spanning three hierarchical capability buckets whose tasks require executable outputs verified by deterministic checkers or simulator-based evaluators. On SCBench, three frontier models exhibit monotonically decreasing accuracy up the capability ladder. Sweeping output-token caps shows that accuracy gains concentrate at low budgets and saturate quickly, and failures are dominated by locally plausible geometry that breaks global constraints. We release the task generators, verifiers, and visualisation tooling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。