arXiv:2609.03681cs.RO2026-09

用世界模型智能调度想象,让视觉语言动作模型高效训练。

WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models

论文配图:WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models
图 1 · 摘自论文原文
  • 根据任务阶段选择性使用世界模型想象,避免无效计算。
  • 多视角有限滚动评估未来,减少误差累积,提升学习信号可信度。
  • 实测节省80%算力,真实场景下泛化与鲁棒性显著提升。

后训练视觉-语言-动作(VLA)策略通常依赖昂贵的专家示范监督微调,或代价高昂且可能不稳定的现实世界探索强化学习。世界模型通过预测未来行为结果提供替代方案,但有效后训练不仅需要精准预测:想象力需在合适时机使用、控制在可靠时间范围内,并转化为可信的策略监督。在机器人操作任务中,想象的价值随执行阶段变化显著,而过长的轨迹推演会累积预测误差,引入不可靠的学习信号。本文提出WISE(世界模型引导的想象调度框架),统一协调策略优化过程中世界模型的使用时机与方式。WISE在交互相关状态时选择性触发想象,进行有限多视角滚动推演,利用进展与完成度信号评估候选未来,并以相对结果修正来自真实交互的动作。大量实验表明,在$\pi_0$和$\pi_{0.5}$设置下,各项操作任务均有稳定提升,相比全程想象减少约80%的GPU计算时间。真实世界验证显示,在多种分布偏移下,系统鲁棒性与泛化能力显著增强。

原文摘要 · Abstract (English)

Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinforcement learning with expensive and potentially unstable real-world exploration. World models offer a promising alternative by evaluating candidate behaviors through imagined futures, yet effective post-training requires more than accurate prediction: imagination must be scheduled where it is useful, bounded within reliable horizons, and translated into trustworthy policy supervision. In robotic manipulation, the value of imagination varies substantially across execution stages, while extended rollouts can accumulate prediction errors and introduce unreliable learning signals. We introduce WISE (World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models), a unified framework that coordinates when and how world-model imagination is used during policy refinement. WISE selectively invokes imagination at interaction-relevant states, performs bounded multi-view rollouts, evaluates candidate futures using progress and completion signals, and uses their relative outcomes to refine actions generated from real interaction contexts. Extensive experiments with both $\pi_0$ and $\pi_{0.5}$ demonstrate consistent improvements across diverse manipulation tasks while reducing GPU computation time by approximately 80% compared with full imagination. Real-world evaluations further show substantial gains in robustness and generalization under diverse real-world distribution shifts.

视觉语言动作世界模型智能调度高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。