构建分层基准测试STARK,评估大模型在时空推理中的能力与局限。
Benchmarking Spatiotemporal Reasoning in LLMs and Reasoning Models: Capabilities and Challenges
- 提出三级推理评估框架:状态估计、时空关系推断、融合世界知识的推理
- 8个LLM和3个LRM在14,552个任务中表现,大模型几何推理能力弱
- 大型推理模型(如o3)整体领先,尤其在复杂任务中表现突出
时空推理在信息物理系统(CPS)中至关重要。尽管大语言模型(LLMs)和大推理模型(LRMs)取得进展,其对复杂时空信号的推理能力仍待深入探索。本文提出分层时空推理基准测试STARK,系统评估模型在三个层次上的表现:状态估计(如场变量预测、事件定位与追踪)、基于状态的时空推理(如推断时空关系)、以及融合上下文与领域知识的世界知识推理(如意图预测、地标感知导航)。我们构建了26个不同模态的时空任务,包含14,552个挑战,模型可通过直接回答或使用Python代码解释器作答。评估3个LRM和8个LLM发现,LLMs在需要几何推理的任务(如多边定位或多点三角测量)中表现有限,且复杂度越高越差。令人意外的是,LRMs在各类难度任务中均表现出色,常优于甚至超越传统基于原理的方法。在需世界知识的任务中,LLMs与LRMs差距缩小,部分LLMs甚至超越后者。但总体而言,大型推理模型o3在所有任务中保持领先,主要归因于其更大的规模。STARK为未来智能CPS的模型架构与推理范式创新提供了结构化框架,有助于识别当前模型在时空推理中的瓶颈。
原文摘要 · Abstract (English)
Spatiotemporal reasoning plays a key role in Cyber-Physical Systems (CPS). Despite advances in Large Language Models (LLMs) and Large Reasoning Models (LRMs), their capacity to reason about complex spatiotemporal signals remains underexplored. This paper proposes a hierarchical SpatioTemporal reAsoning benchmaRK, STARK, to systematically evaluate LLMs across three levels of reasoning complexity: state estimation (e.g., predicting field variables, localizing and tracking events in space and time), spatiotemporal reasoning over states (e.g., inferring spatial-temporal relationships), and world-knowledge-aware reasoning that integrates contextual and domain knowledge (e.g., intent prediction, landmark-aware navigation). We curate 26 distinct spatiotemporal tasks with diverse sensor modalities, comprising 14,552 challenges where models answer directly or by Python Code Interpreter. Evaluating 3 LRMs and 8 LLMs, we find LLMs achieve limited success in tasks requiring geometric reasoning (e.g., multilateration or triangulation), particularly as complexity increases. Surprisingly, LRMs show robust performance across tasks with various levels of difficulty, often competing or surpassing traditional first-principle-based methods. Our results show that in reasoning tasks requiring world knowledge, the performance gap between LLMs and LRMs narrows, with some LLMs even surpassing LRMs. However, the LRM o3 model continues to achieve leading performance across all evaluated tasks, a result attributed primarily to the larger size of the reasoning models. STARK motivates future innovations in model architectures and reasoning paradigms for intelligent CPS by providing a structured framework to identify limitations in the spatiotemporal reasoning of LLMs and LRMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。