arXiv:2607.26712cs.RO2026-07被引 1

提出可区分不同动作的潜空间世界模型,提升长时序规划稳定性。

ActSWM: Action-Sensitive World Models for Long-Horizon Planning in Open-World Games

论文配图:ActSWM: Action-Sensitive World Models for Long-Horizon Planning in Open-World Games
图 1 · 摘自论文原文
  • 基于动作可区分性约束潜空间演化,防止未来状态趋同
  • 在Minecraft中实现90%以上的任务成功率,优于基线模型
  • 适用于游戏智能体长期决策与离线动作恢复

潜空间世界模型通过在潜空间优化未来控制序列并以滚动时域方式重规划,支持高效的模型预测控制。然而,现有潜空间预测器常缺乏稳定的长时序回放能力,且预测精度不足以保证回放对所规划动作保持响应。我们识别出一种称为上下文坍缩(Context Collapse)的失效模式:自回归潜空间预测器虽保持未来状态间高相似度,但在不同动作序列下生成几乎无法区分的未来。为解决此问题,我们提出ActSWM,一种基于迁移分离原则的动作敏感潜空间世界模型:一个可用于规划的潜动态模型应保持不同动作未来可区分,并使每个局部转移对应的动作可恢复。在此原则下,动作敏感性作为潜回放中的约束而非辅助目标,促使预测未来在长时序中保留动作依赖差异。在步移分析、闭环Minecraft规划和跨游戏局部动作恢复实验中,ActSWM保持的行动依赖回放差距显著大于基线模型,在长时序交互设置中提升任务成功率,并实现从离线游戏视频中基于世界模型的动作恢复。

原文摘要 · Abstract (English)

Latent world models support efficient model-predictive control by optimizing future control sequences in latent space and replanning in a receding-horizon manner. However, existing latent predictors often lack stable long-horizon rollout ability, and prediction accuracy alone does not ensure that rollouts remain responsive to the actions being planned. We identify Context Collapse, a failure mode in which autoregressive latent predictors maintain high similarity to future states while producing nearly indistinguishable futures under different action sequences. To address this issue, we propose ActSWM, an action-sensitive latent world model grounded in a transition-separation principle: a planning-useful latent dynamics model should keep alternative-action futures distinguishable and make the action associated with each local transition recoverable. Under this principle, action sensitivity is enforced as a constraint on latent rollouts rather than treated only as an auxiliary prediction target, encouraging predicted futures to preserve action-dependent differences over long horizons. Across step-drift analysis, closed-loop Minecraft planning, and cross-game local action recovery, ActSWM preserves larger action-dependent rollout gaps than existing baselines, improves task success in long-horizon interactive settings, and enables world-model-based action recovery from offline gameplay videos.

世界模型长时规划动作敏感Minecraft

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。