arXiv:2511.02627cs.AI2025-11

构建可分解的空间推理数据集,精准测试大模型的组合推理能力。

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning

  • 通过程序化生成,独立控制推理深度、语言变化等成分性因素。
  • 超500万数据点,验证显示大模型在深层推理与新语言元素上表现差。
  • 适合研究大模型空间推理与组合泛化能力的学者使用。

我们提出 DecompSR,一个大型基准数据集(超过500万样本)及生成框架,用于分析组合式空间推理能力。DecompSR 的生成机制可独立调节多种组合性特征:生产力(推理深度)、可替换性(实体与语言变体)、过度泛化(输入顺序、干扰项)和系统性(新语言元素)。该数据集通过程序化构建实现构造正确性,并经符号求解器独立验证以确保准确性。我们在多种大语言模型(LLMs)上对 DecompSR 进行全面评估,结果表明:尽管模型对语言变化较鲁棒,但在生产性推理和系统性泛化方面表现不佳。DecompSR 提供了一个可证明正确的严格基准,支持对大模型组合推理能力进行细致、精准的探查。

原文摘要 · Abstract (English)

We introduce DecompSR, decomposed spatial reasoning, a large benchmark dataset (over 5m datapoints) and generation framework designed to analyse compositional spatial reasoning ability. The generation of DecompSR allows users to independently vary several aspects of compositionality, namely: productivity (reasoning depth), substitutivity (entity and linguistic variability), overgeneralisation (input order, distractors) and systematicity (novel linguistic elements). DecompSR is built procedurally in a manner which makes it is correct by construction, which is independently verified using a symbolic solver to guarantee the correctness of the dataset. DecompSR is comprehensively benchmarked across a host of Large Language Models (LLMs) where we show that LLMs struggle with productive and systematic generalisation in spatial reasoning tasks whereas they are more robust to linguistic variation. DecompSR provides a provably correct and rigorous benchmarking dataset with a novel ability to independently vary the degrees of several key aspects of compositionality, allowing for robust and fine-grained probing of the compositional reasoning abilities of LLMs.

空间推理组合性大模型评测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。