arXiv:2604.26848cs.RO2026-04被引 8

让机器人更精准地预判动作空间时间变化,提升复杂操作成功率。

STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation

论文配图:STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation
图 1 · 摘自论文原文
  • 用统一扩散过程同时去噪未来动作和空间时间特征,实现动作与预测对齐。
  • 在50个双臂任务中,模拟成功率93.82%(洁净环境),真实世界提升至70.8%。
  • 适合需要精细空间时间协调的机器人抓取、装配等复杂操作任务。

机器人操作需推理未来的时空交互与几何约束,但现有视觉-语言-动作(VLA)策略常使预测表示与动作执行耦合松散,导致精确时空协调任务失败。本文提出STARRY,一种基于世界模型的动作生成策略,通过统一扩散过程联合去噪未来时空潜在变量与动作,实现预测与动作的对齐。为连接2D视觉标记与3D度量控制,STARRY引入几何感知选择性注意力调制(GASAM),将预测深度与末端执行器几何转换为标记对齐权重,用于选择性动作注意力调制。在RoboTwin 2.0上,STARRY在50个双臂任务中,洁净与随机设置下平均成功率达93.82% / 93.30%。真实世界实验表明,相比π₀.₅,平均成功率从42.5%提升至70.8%。结果验证了以动作为中心的时空世界建模在高要求空间-时间操作中的有效性。

原文摘要 · Abstract (English)

Robotic manipulation requires reasoning about future spatial-temporal interactions and geometric constraints, yet existing Vision-Language-Action (VLA) policies often leave predictive representation weakly coupled with action execution, causing failures in tasks requiring precise spatial-temporal coordination. We propose STARRY, a world-model-enhanced action-generation policy that aligns spatial-temporal prediction and action generation by jointly denoising future spatial-temporal latents and actions through a unified diffusion process. To bridge 2D visual tokens and 3D metric control, STARRY introduces Geometry-Aware Selective Attention Modulation (GASAM), which converts predicted depth and end-effector geometry into token-aligned weights for selective action-attention modulation. On RoboTwin 2.0, STARRY achieves 93.82% / 93.30% average success under Clean and Randomized settings across 50 bimanual tasks. Real-world experiments show that STARRY improves average success from 42.5% to 70.8% compared with $π_{0.5}$. These results demonstrate the effectiveness of action-centric spatial-temporal world modeling for spatially and temporally demanding robotic manipulation.

机器人操作时空建模扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。