arXiv:2512.18477cs.ROcs.CV2025-12被引 6

用搜索引导生成模型,让机器人更聪明地规划复杂操作。

STORM: Search-Guided Generative World Models for Robotic Manipulation

  • 用扩散模型生成动作候选,视频世界模型模拟视觉与奖励结果
  • 在仿真中预测未来,提升任务成功率至51.0%,减少75%以上视频误差
  • 适合需要长程规划与失败恢复的机器人任务场景

我们提出STORM(搜索引导的生成式世界模型),一种用于机器人操作中的时空推理新框架,统一了基于扩散的动作生成、条件视频预测和搜索规划。与依赖抽象潜态动力学或将推理交给语言模块的现有视觉-语言-动作(VLA)模型不同,STORM将规划建立在显式的视觉回放之上,实现可解释且具前瞻性的决策。基于扩散的VLA策略生成多样化动作候选,生成式视频世界模型模拟其视觉与奖励结果,蒙特卡洛树搜索(MCTS)通过前瞻评估筛选并优化计划。在SimplerEnv操作基准上的实验表明,STORM实现了51.0%的平均成功率,创下新纪录,显著优于CogACT等强基线。奖励增强的视频预测大幅提升了时空保真度与任务相关性,使弗雷谢尔视频距离降低超过75%。此外,STORM展现出强大的重规划与故障恢复能力,凸显搜索引导生成式世界模型在长时序机器人操作中的优势。

原文摘要 · Abstract (English)

We present STORM (Search-Guided Generative World Models), a novel framework for spatio-temporal reasoning in robotic manipulation that unifies diffusion-based action generation, conditional video prediction, and search-based planning. Unlike prior Vision-Language-Action (VLA) models that rely on abstract latent dynamics or delegate reasoning to language components, STORM grounds planning in explicit visual rollouts, enabling interpretable and foresight-driven decision-making. A diffusion-based VLA policy proposes diverse candidate actions, a generative video world model simulates their visual and reward outcomes, and Monte Carlo Tree Search (MCTS) selectively refines plans through lookahead evaluation. Experiments on the SimplerEnv manipulation benchmark demonstrate that STORM achieves a new state-of-the-art average success rate of 51.0 percent, outperforming strong baselines such as CogACT. Reward-augmented video prediction substantially improves spatio-temporal fidelity and task relevance, reducing Frechet Video Distance by over 75 percent. Moreover, STORM exhibits robust re-planning and failure recovery behavior, highlighting the advantages of search-guided generative world models for long-horizon robotic manipulation.

机器人操作生成模型搜索规划视频预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。