arXiv:2607.10362cs.LG2026-07被引 2

提出新目标函数,让潜空间模型更准确预测规划路径的代价。

A Control Theory of Predictability in Latent World Models

  • 用规划路径的实际代价与预测代价的差距作为训练目标
  • 单步预测误差与控制成功率几乎无关,但新指标高度相关
  • 适合研究模型预测控制与强化学习中可预测性问题的人

潜空间世界模型通过学习表示来预测未来状态,并在规划器中用于选择动作。当前做法以测试集上的预测误差(单步或多步滚动损失)作为训练和模型选择目标,假设更低的误差带来更好的控制性能。我们指出这一假设在结构上不可靠:规划器并不在训练分布上查询模型,而是在其候选动作所到达的状态上查询,这些状态通常脱离数据流形。因此,对数据分布平均的误差无法单独决定控制表现。我们重新定义目标为规划器最终选定计划的真实代价与预测代价之间的差异,并证明规划器的次优性由该差异的两倍所限制;而数据平均预测误差既不能界定也不能跟踪此次优性。在线性控制假设下,该差异分解为两项:一是流形内的残差误差,预测与真实动态一致,由潜在转移算子的非正规性定价;二是流形外的发散项,动作将状态带离流形后两者动态分离,是主导项且无法被数据平均误差约束。合成算子验证了定价公式,潜空间模型预测控制实验验证了解耦:在不同随机种子下,单步验证误差与控制成功几乎不相关,而规划器可达测度上的保真度得分则能有效追踪控制表现。

原文摘要 · Abstract (English)

Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward. Current practice adopts the prediction error, the single- or multi-step rollout loss on held-out data, as the training and model-selection objective, on the assumption that a lower prediction error yields better control. We show that this assumption is unreliable for a structural reason: a planner does not query the model on the training distribution but on the states that its candidate actions reach, which generally leave the data manifold, so an error averaged over the data cannot by itself govern control. We therefore reframe the objective as the discrepancy between the predicted and the true plan-cost at the plan the planner commits to, and prove that the planner's suboptimality is bounded by twice this discrepancy, whereas the data-averaged prediction error neither bounds nor tracks it. Under a linear-control premise the discrepancy separates into two terms. The first is a small on-manifold residual, on which the predicted and true dynamics agree and which a spectral tax prices through the non-normality of the latent transition operator. The second is an off-manifold divergence, on which an action carries the state off the manifold and the two dynamics diverge; this divergence is the binding term and is bounded by no data-averaged error. Synthetic operators confirm the pricing formulas, and latent model-predictive control experiments confirm the decoupling: across seeds, the single-step validation error is essentially uncorrelated with control success, whereas a fidelity score on the planner-reachable measure tracks it.

世界模型控制理论可预测性规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。