世界模型的干预真实性需直接检测,否则易被奖励拟合误导。
The Intervention Gap in Latent World Models

- 通过开环模拟测试模型对任务变量的干预效果是否真实
- 模型奖励误差小但实际干预效果差,且任务训练反使表现更糟
- 适合关注模型可解释性与真实决策能力的研究者
计划阶段的干预保真度是可度量的模型属性:模型自身的开环转移是否与环境干预一致。在测试中,该属性既不反映在奖励拟合上,也无法通过任务锚定训练保证。在多个释放的TD-MPC2检查点中,随着操作误差诊断指标上升,平均回报下降,而奖励预测误差保持小且稳定;无任务信号的自监督世界模型在相同任务上表现远优于任务锚定模型。捕获门控匹配干预审计定位失败原因:在Cheetah任务中,三个LeWorldModel检查点能捕捉当前查询并解码真实干预效应,但其五步想象效应劣于预测无效应,甚至低于环境终点的基准。失败源于任务方向旋转与增益过大,而非特征坍塌。此问题具有条件性:五个PreJEPA种子中,部分保留基准缺陷;Finger Spin实验显示该缺陷扩展至非运动任务,且不同种子严重程度各异;共享银行效应几何结构依赖候选与支持。此外,在DreamerV3中,后验分布承载当前查询,而非采样值;集成分歧仅在训练支持附近有效;冻结的支持感知评分在两种迁移方向均降低泛化误差排序,而原生分歧仍具信息量。结论:必须直接、以捕获为先,在模型本机接口上审计干预保真度。
原文摘要 · Abstract (English)
Planning-time intervention fidelity is a distinct, measurable property of a learned world model: whether the model's own open-loop transitions move task variables the way matched environment interventions do. In the settings we test, it is neither revealed by reward fit nor ensured by task-anchored training. Across released TD-MPC2 checkpoint sizes, episode return falls as an operator-error diagnostic on task observables grows, while reward-prediction error stays small and nearly flat, and a self-supervised world model trained without task signal preserves the same operator substantially better than a task-anchored model on the shared task. A capture-gated matched-intervention audit then localizes what fails. On Cheetah, three LeWorldModel checkpoints capture the current task query and support decodable real intervention effects; however, their imagined five-step effects are worse than predicting no effect and worse than an environment-endpoint oracle. The failure is task-direction rotation with excess gain, not feature collapse. This severe pattern is conditional: five PreJEPA seeds retain an oracle-relative deficit without it, Finger Spin experiments extend the deficit beyond locomotion with heterogeneous severity across seeds, and shared-bank effect geometry is both candidate- and support-dependent. We also test practice-side questions. In DreamerV3 the posterior distribution, not its sample, carries the current query; ensemble disagreement ranks error only near training support; and a frozen support-aware score degrades held-out error ranking in both tested transfer directions while native disagreement remains informative in both. We conclude that intervention fidelity must be audited directly, capture-first, on the model's native interface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。