arXiv:2609.06578cs.CV2026-09

让机器人模型根据执行进度动态调整想象未来,提升动作决策可靠性。

Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models

论文配图:Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models
图 1 · 摘自论文原文
  • 用执行进度作为中间表示,动态调节对未来的想象使用
  • 在多个任务上相比基线提升12.3%成功率,尤其在长序列任务中优势明显
  • 适合需要连续推理的机器人控制场景,如家庭服务或工业操作

世界动作模型(WAMs)通过引入未来视觉动态来扩展视觉-语言-动作(VLA)模型。然而,现有WAMs对想象未来的利用缺乏对执行进展的适应性,可能引入干扰或不可靠的预测线索。这一局限源于未来效用的两种非均匀性:(i)跨执行阶段间,控制需求变化导致未来效用波动;(ii)同一阶段内,不同未来潜变量相关性差异显著。为此,我们提出ProWAM,一种基于执行进度的自适应世界动作模型。ProWAM包含两个紧密耦合组件:(1) 自监督双时序进度编码器(SS-DTPE),通过建模短期动作-观测交互与长期递归进度聚合,捕捉近期反馈与累积任务历史;(2) 分层进度条件想象调制(HPIM),基于SS-DTPE输出的进度表示,实现跨阶段全局调制与同阶段内部相关性区分。大量实验表明,ProWAM在多个基准上持续优于强基线,平均提升12.3%成功率。

原文摘要 · Abstract (English)

World Action Models (WAMs) extend Vision-Language-Action (VLA) models by incorporating future visual dynamics into action generation. However, existing WAMs often utilize imagined futures with limited adaptation to evolving execution progress, potentially introducing distracting or unreliable predictive cues. This limitation arises from two empirically identified forms of non-uniformity in future utility: (i) at the inter-progress level, the utility of imagined futures varies across execution stages as control demands change; and (ii) at the intra-progress level, individual future latents exhibit heterogeneous relevance within the same progress state. To address these limitations, we propose ProWAM, a Progress-Conditioned World Action Model that introduces execution progress as an explicit intermediate representation for adaptive imagination utilization. ProWAM comprises two tightly coupled components: (1) To obtain a reliable representation of execution progress, we propose the Self-Supervised Dual-Temporal Progress Encoder (SS-DTPE). SS-DTPE couples short-term action-observation interaction modeling with long-term recurrent progress aggregation to capture recent execution feedback and accumulated task history. (2) Conditioned on the progress representation from SS-DTPE, we propose the Hierarchical Progress-Conditioned Imagination Modulation (HPIM) to adapt imagination utilization to execution progress. HPIM operates at two complementary levels: an inter-progress global modulation mechanism adapts future utilization across execution stages, while an intra-progress relevance mechanism differentiates individual future latents within each progress state. Extensive experiments demonstrate consistent gains over strong VLA and WAM baselines.

机器人控制动作规划想象生成进度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。