用3D-4D统一表征提升机器人在复杂环境中的时空理解能力。
ST-VLA: Enabling 4D-Aware Spatiotemporal Understanding for General Robot Manipulation
- 构建3D-4D联合表征,实现语义与动作的稳定衔接。
- 在RLBench上零样本成功率提升44.6%,真实场景提升30.3%。
- 适合需要长时序推理与3D空间感知的机器人任务研究者。
开放世界中的机器人操作需跨语义、几何与长时序动作动态进行推理。现有分层视觉-语言-动作(VLA)框架多采用2D表示连接高层推理与底层控制,但缺乏深度感知与时间一致性,限制了在复杂3D场景中的鲁棒性。本文提出ST-VLA,一种基于统一3D-4D表示的分层VLA框架,将2D指令转换为3D轨迹,并生成平滑的空间掩码以捕捉4D时空上下文,为语义推理与连续控制提供稳定接口。为有效学习此类表征,我们构建了大规模人类操作数据集ST-Human,包含14项任务与30万条轨迹,通过半自动化流程标注2D、3D及4D监督信号。基于此,训练出可生成空间对齐且时间连贯的3D表示的时空视觉-语言模型ST-VLM,其平滑空间掩码聚焦任务相关几何结构,稳定潜在表示,支持在线重规划与长时序推理。在RLBench和真实世界任务上的实验表明,该方法显著优于现有基线,零样本成功率分别提升44.6%和30.3%。结果表明,将时空推理任务交由具备统一3D-4D表示的VLM可大幅提高开放世界机器人操作的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Robotic manipulation in open-world environments requires reasoning across semantics, geometry, and long-horizon action dynamics. Existing hierarchical Vision-Language-Action (VLA) frameworks typically use 2D representations to connect high-level reasoning with low-level control, but lack depth awareness and temporal consistency, limiting robustness in complex 3D scenes. We propose ST-VLA, a hierarchical VLA framework using a unified 3D-4D representation to bridge perception and action. ST-VLA converts 2D guidance into 3D trajectories and generates smooth spatial masks that capture 4D spatio-temporal context, providing a stable interface between semantic reasoning and continuous control. To enable effective learning of such representations, we introduce ST-Human, a large-scale human manipulation dataset with 14 tasks and 300k episodes, annotated with 2D, 3D, and 4D supervision via a semi-automated pipeline. Using ST-Human, we train ST-VLM, a spatio-temporal vision-language model that generates spatially grounded and temporally coherent 3D representations to guide policy execution. The smooth spatial masks focus on task-relevant geometry and stabilize latent representations, enabling online replanning and long-horizon reasoning. Experiments on RLBench and real-world manipulation tasks show that \method significantly outperforms state-of-the-art baselines, improving zero-shot success rates by 44.6% and 30.3%. These results demonstrate that offloading spatio-temporal reasoning to VLMs with unified 3D-4D representations substantially improves robustness and generalization for open-world robotic manipulation. Project website: https://oucx117.github.io/ST-VLA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。