arXiv:2605.29864cs.RO2026-05

用大模型生成未来视频,帮机器人在不确定环境中更好规划动作。

LLM-Guided Future Hypotheses for Horizon-Aware Exploration in Multi-Step Robot Manipulation

论文配图:LLM-Guided Future Hypotheses for Horizon-Aware Exploration in Multi-Step Robot Manipulation
图 1 · 摘自论文原文
  • 用大模型+数字孪生+扩散模型生成短时未来视频作为动作先验
  • 未来视频条件下的策略比无未来信息提升性能,错误未来则拖后腿
  • 适合做多步操作的机器人控制,尤其在环境变化不确定时

多步机器人操作需在场景演化不确定性下进行探索与策略调整。本文研究短时、任务一致的未来视频能否为控制和强化学习微调提供有效结构化先验。提出未来经验条件化(FEC)框架,将闭环策略基于未来视频的隐式表征进行条件化。在仿真中,未来片段通过三阶段生成:基于任务本体的大模型推理、无需机器人的数字孪生动作推演、无需分割的掩码自由视频扩散模型合成。以行为克隆(BC)及BC+RL为主实例化该接口,在RoboCasa和CALVIN数据集上对比了无未来(NoFuture)、真实未来(GTFuture)、生成未来(GenFuture)和错误未来(WrongFuture)四种设置。结果显示,有未来条件的策略优于无未来,错误未来导致性能下降;其中BC+RL整体表现最佳。8个CALVIN任务的平均学习曲线分析表明,GTFuture收敛最快,GenFuture早期提升更早且达到更高水平,而WrongFuture始终无效。结果表明,短时未来视频可作为探索与策略适应的有用结构化先验。

原文摘要 · Abstract (English)

Multi-step robot manipulation requires acting under uncertainty about how the scene will evolve, making exploration and policy adaptation challenging. We study whether short-horizon, task-consistent future videos can provide useful structured priors for control and reinforcement-learning fine-tuning. We formalize this idea through Future-Experience Conditioning (FEC), a simple interface that conditions closed-loop policies on a latent representation of a short future video. In our simulation setup, future clips are generated in three stages, an LLM reasoner operating over a task ontology initialized from the current scene state, a robot-free digital-twin rollout of the intended object motion, and a mask-free video diffusion model that synthesizes a robot-consistent future clip without requiring segmentation at inference. We instantiate this future-conditioning interface primarily with BC and BC+RL, and compare against a future-conditioned Streaming Flow Policy (SFP) baseline on RoboCasa and CALVIN under NoFuture, GTFuture, GenFuture, and WrongFuture. Generated futures improve performance over no-future conditioning, while mismatched futures degrade it, and our BC+RL instantiation achieves the strongest overall results. An average BC+RL learning-curve analysis across 8 CALVIN tasks further shows that GTFuture improves fastest, GenFuture improves earlier and to a higher level than NoFuture, and WrongFuture remains at zero throughout training. These results suggest that short-horizon future videos can serve as useful structured priors for exploration and policy adaptation under imperfect future predictions. https://enact2026.github.io/

机器人操作大模型未来预测强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。