构建时空道路图像数据集,助力自动驾驶模型理解环境动态变化。
TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy
- 将360度全景图构建成时空互联的观测节点结构
- 在STRIDE数据集上训练出性能优异的生成世界模型
- 适合研究自动驾驶、具身智能与时空建模的学者
世界模型旨在模拟环境并实现有效代理行为。然而,真实世界环境在空间和时间维度上均呈现动态变化,建模面临独特挑战。为此,我们提出用于探索的时空道路图像数据集STRIDE,将360度全景影像转换为丰富互联的观察、状态与动作节点。基于该结构,可同时建模自车视角、位置坐标与运动指令在时空维度上的关系。我们通过TARDIS——一个基于Transformer的生成式世界模型,在统一自回归框架下整合时空动态,并在STRIDE上进行训练。结果表明,该模型在可控照片级图像生成、指令遵循、自主自我控制及最先进地理定位任务中表现稳健。这些成果展示了通向具备空间与时间理解能力的通用智能体的可行路径,显著提升其具身推理能力。训练代码、数据集及模型检查点已公开于https://huggingface.co/datasets/Tera-AI/STRIDE。
原文摘要 · Abstract (English)
World models aim to simulate environments and enable effective agent behavior. However, modeling real-world environments presents unique challenges as they dynamically change across both space and, crucially, time. To capture these composed dynamics, we introduce a Spatio-Temporal Road Image Dataset for Exploration (STRIDE) permuting 360-degree panoramic imagery into rich interconnected observation, state and action nodes. Leveraging this structure, we can simultaneously model the relationship between egocentric views, positional coordinates, and movement commands across both space and time. We benchmark this dataset via TARDIS, a transformer-based generative world model that integrates spatial and temporal dynamics through a unified autoregressive framework trained on STRIDE. We demonstrate robust performance across a range of agentic tasks such as controllable photorealistic image synthesis, instruction following, autonomous self-control, and state-of-the-art georeferencing. These results suggest a promising direction towards sophisticated generalist agents--capable of understanding and manipulating the spatial and temporal aspects of their material environments--with enhanced embodied reasoning capabilities. Training code, datasets, and model checkpoints are made available at https://huggingface.co/datasets/Tera-AI/STRIDE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。