探究潜在世界模型能识别哪些物理参数,发现预测目标决定信息留存。
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

- 通过可控干预实验,检验物理参数在潜在表示中的可辨识性。
- 触觉预测使刚度信息进入潜空间(R²=0.50),视觉仅预测则无效。
- 预测目标结构决定信息获取,数据量提升仅增强已学参数性能。
潜在世界模型的核心假设是:未来预测迫使表示内化环境物理规律。训练后的潜在表示究竟包含哪些物理量?由何决定?我们通过在POKEWORLD环境中进行受控干预来回答。该环境中的视觉相同物体隐藏质量、阻力和接触刚度参数。采用证书门控协议,先验证各参数是否可从原始观测中恢复,再测量其是否进入潜在表示,确保零结果归因于目标而非环境。所得可辨识图揭示两个组织机制与一个前沿:输入限制可知晓内容,预测目标决定保留内容。刚度仅在预测触觉时进入潜空间(R²=0.50),而仅融合信号时为-0.02;单步预测下,仅视觉潜空间丢弃完全可见的物体状态。阻力处于边界:可恢复性证书为0.89,但在所有确定性预测目标下均停滞于0.13,而同一主干的监督头可达0.45。读出缓慢且依赖比值的参数无法被这些目标获取。在RH20T上,跨缩放曲线的输入-目标因子实验在两机器人、4,258个轨迹中复现上述机制:缺失信息或预测压力的臂在五倍数据范围内保持平坦,仅全模态目标能超越持续基线,且外推收益随规模增长。目标结构决定了潜空间所获物理参数,额外数据仅提升已有参数的性能。
原文摘要 · Abstract (English)
A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this? We answer with controlled interventions in POKEWORLD, an interactive environment whose visually identical objects hide mass, drag, and contact stiffness. A certificate-gated protocol first certifies each parameter as recoverable from raw observations, then measures whether it enters the latent, so a null result can be attributed to the objective rather than to the environment. The resulting identifiability map has two organizing mechanisms and one frontier. Inputs limit what can be known, while prediction targets decide what is retained. Stiffness enters the latent only when touch is forecast ($R^2=0.50$, compared with $-0.02$ when the same signal is merely fused into the input), and under single-step prediction a vision-only latent discards even perfectly visible object state. Drag marks the frontier. It carries a recoverability certificate of 0.89 yet plateaus near 0.13 under every deterministic prediction objective we test, while a supervised head on the same trunk reaches 0.45. Parameters whose readout is slow and ratio-type under the sensed coordinates fall outside what these objectives acquire. On RH20T, an input-target factorial across scaling curves reproduces both mechanisms across two robots and 4,258 episodes. Every arm missing information or prediction pressure stays flat over a fivefold data range, and only the full multimodal objective forecasts force beyond a persistence baseline, with held-out gains that grow with scale. Objective structure determines which physical parameters a latent acquires, and additional data improves only the parameters it already acquires.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。