arXiv:2605.18380cs.AI2026-05

新基准测试大模型在空间时间推理上的能力,发现不同模型表现差异大。

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi

论文配图:QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi
图 1 · 摘自论文原文
  • 构建涵盖多种时空逻辑的综合推理测试集
  • 模型对点代数易答,区域连接22种关系最难
  • 适合研究逻辑推理与认知建模的学者参考

我们提出一个面向大语言模型的定性空间与时间推理(QSTR)评估基准。问题涉及组合推理(使用组合表)、逆关系和概念邻域(CN),覆盖点代数(PA)、Allen区间代数、区间与持续时间(INDU)、区域连接计算(RCC-5、RCC-8、RCC-22)、九交集模型、基数方向计算和STAR。首次发布RCC-22的概念邻域。基准系统性地改变问题呈现方式,包括前缀/中缀、词语/符号/虚构术语及图示描述。报告了前沿模型的表现结果:所有模型均优于随机猜测,但无一能全部正确回答。性能随逻辑体系差异显著,其中点代数最简单,RCC-22最困难。我们开源该基准及评测结果,以推动对大模型定性时空推理能力的研究。

原文摘要 · Abstract (English)

We introduce an extensive qualitative spatial and temporal reasoning (QSTR) benchmark for evaluating large language models (LLMs). We pose questions concerning compositional reasoning (using composition tables, CT), converse relations, and conceptual neighbourhoods (CN) for QSTR calculi, Point Algebra (PA), Allen's Interval Algebra, Interval and Duration (INDU), Region Connection Calculus (RCC-5, RCC-8, and RCC-22), the nine intersection model, cardinal direction calculus, and STAR. The RCC-22 CN is published here for the first time. An extended benchmark systematically varies question presentation including prefix/infix, words/symbols/nonce terms and schematic descriptions for selected calculi. We report results for contemporary frontier models. All models tested perform better than guessing but none can consistently answer all questions correctly. Performance varies sharply by calculus, with PA being the most straightforward, and RCC-22 the most difficult. We release the benchmark, and our results under an open licence to facilitate further assessment of qualitative spatio/temporal reasoning in LLMs.

空间推理时间逻辑大模型评测认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。