让潜空间距离更真实反映任务进展,提升规划效率
Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

- 用排序一致性指标诊断潜空间规划的偏差问题
- 引入新目标函数后收敛更快,成功率达87%以上
- 适合做基于潜空间的强化学习与机器人规划研究
JEPA风格的潜空间世界模型可使用目标潜变量间的欧氏距离作为模型预测控制(MPC)的代价。然而,强任务变量解码并不保证该代价能正确排序候选动作序列的真实任务进展。我们称此性质为“决策-度量对齐”。提出Plan-Real Spearman和CEM-stage Spearman两个指标,分别衡量随机规划和交叉熵法(CEM)优化过程中潜空间排名与真实任务进展的一致性。分析表明,编码器失真、终端滚动误差和候选间隔是影响对齐的关键因素。基于实证对齐差距,提出DA-LeWM,在LeWM基础上加入逆动力学和示范条件下的目标-动作头。在所有实验中,DA-LeWM加速收敛并实现更高在线成功率(最高达87%),而探针分数保持相似。结果表明,动作条件目标能有效改善欧氏代价与实际任务进展之间的几何关系。
原文摘要 · Abstract (English)
JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). Strong decoding of task variables, however, does not guarantee that this particular cost ranks candidate action sequences by real task progress. We call the latter property \emph{decision-metric alignment}. We introduce Plan-Real Spearman, which measures latent--real rank agreement on random plans, and CEM-stage Spearman, which measures the same agreement as cross-entropy-method (CEM) search concentrates its proposal. We analyze sufficient conditions under which latent distance preserves real-cost rankings, identifying encoder distortion, terminal rollout error, and candidate margins as the controlling quantities. Guided by the observed empirical alignment gap, DA-LeWM augments LeWM with inverse-dynamics and demonstration-conditioned goal-action heads. Across all our experiments, DA-LeWM accelerates convergence and achieves higher online success than LeWM, while probe scores remain similar. These results show that action-conditioned objectives improve the geometry used by Euclidean-cost, CEM-based latent MPC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。