arXiv:2602.12218cs.LGcs.AI2026-02被引 2

用非侵入评估法发现:模型真懂物理,但传统测试会骗人。

The Observer Effect in World Models: Invasive Adaptation Corrupts Latent Physics

  • 用冻结表征+低容量探测器,避免干扰学习到的物理结构
  • 在流体与轨道场景中,OOD下仍能线性解码能量和平方反比定律(ρ>0.90)
  • 传统微调评估会破坏物理结构(ρ≈0.05),适合研究世界模型的可信度

判断神经网络是否真正内化物理规律而非依赖统计捷径,尤其在分布外(OOD)情形下仍具挑战。现有评估常通过下游适应(如微调或高容量探测器)测试潜在能力,但此类操作可能改变被测量的表征,从而混淆自监督学习(SSL)中实际学到的内容。本文提出非侵入式评估协议PhyIP,基于线性表征假说,测试冻结表征中物理量是否可线性解码。在流体动力学与轨道力学任务中,当SSL误差较低时,潜在结构具备线性可解码性。PhyIP在OOD测试中成功恢复内能与牛顿平方反比关系(ρ>0.90);而基于适应的评估则使该结构崩溃(ρ≈0.05)。结果表明,适应性评估可能掩盖潜在结构,低容量探测器更适合作为物理世界模型的准确评估工具。

原文摘要 · Abstract (English)

Determining whether neural models internalize physical laws as world models, rather than exploiting statistical shortcuts, remains challenging, especially under out-of-distribution (OOD) shifts. Standard evaluations often test latent capability via downstream adaptation (e.g., fine-tuning or high-capacity probes), but such interventions can change the representations being measured and thus confound what was learned during self-supervised learning (SSL). We propose a non-invasive evaluation protocol, PhyIP. We test whether physical quantities are linearly decodable from frozen representations, motivated by the linear representation hypothesis. Across fluid dynamics and orbital mechanics, we find that when SSL achieves low error, latent structure becomes linearly accessible. PhyIP recovers internal energy and Newtonian inverse-square scaling on OOD tests (e.g., $ρ> 0.90$). In contrast, adaptation-based evaluations can collapse this structure ($ρ\approx 0.05$). These findings suggest that adaptation-based evaluation can obscure latent structures and that low-capacity probes offer a more accurate evaluation of physical world models.

世界模型物理规律评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。