预测目标会忽略不可控但重要的控制特征,仅用少量奖励标签即可修复。
Predictive Objectives Discard Exogenous Control-Relevant Features: A Controlled Mechanistic Study

- 通过可控性与相关性分离实验,发现预测目标偏好可预测性而非控制相关性。
- 六种无奖励目标中,98%以上准确率的控制相关特征均被丢弃。
- 仅需2%奖励标签就能恢复关键特征,适用于多种环境和隐空间维度。
联合嵌入式预测(JEPA)目标通过预测未来隐状态学习表征,但在过程中会丢弃不可控但对控制重要的外源特征,即使这些特征易于编码。这是因为目标优化的是时间可预测性,而非控制相关性。我们设计了一个受控的2×2实验,独立调节特征的可控性和相关性,使用可调预测性旋钮将预测性与控制相关性解耦。对比六种目标:重建、JEPA、动作条件化JEPA、基于可控性的JEPA、随机策略下的逆动力学、奖励引导的JEPA,发现所有无奖励预测目标均使外源控制相关特征的预测准确率接近随机水平(约50%),而奖励引导版本则选择性保留该特征。修复方法标签效率高且鲁棒:仅需2%的奖励标注过渡即可恢复特征,该效果在两种不同表面形式的环境中均成立,且在16到1024的隐空间维度下持续有效。对比双模拟理论预测的潜在几何结构,JEPA学习的隐空间仅实现监督参考模型所达到类间分离的一小部分。
原文摘要 · Abstract (English)
Joint-embedding predictive (JEPA-style) objectives learn representations by predicting future latents. In doing so they can discard features that are exogenous (uncontrollable by the agent) yet control-relevant, even when those features are trivially encodable. This occurs because the objective optimizes temporal predictability rather than control-relevance. We isolate this failure mode in a controlled 2x2 experimental design that varies feature controllability and relevance independently, using a predictability knob that decouples a feature's temporal predictability from its control-relevance. Comparing six objectives: reconstruction, JEPA, action-conditioned JEPA, controllability-based JEPA, inverse dynamics under a random policy, and reward-grounded JEPA, we observe that all evaluated reward-free predictive objectives leave the exogenous control-relevant feature near chance accuracy, while a reward-grounded variant retains it selectively. The remedy is label-efficient and robust: as little as 2% of reward-labeled transitions recovers the feature, the effect holds across two environments with different surface forms, and it persists across latent dimensions from 16 to 1024. Comparing the learned latent geometry against bisimulation theory's prediction, the JEPA latent realizes only a small fraction of the class separation a supervised reference attains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。