无语言监督下,物理交互让世界模型自发学习空间语义结构。
Emergent Semantic Representations in World Models through Physical Interaction without Linguistic Supervision

- 通过随机身体探索训练变分自编码器,潜空间自发形成物理几何结构。
- 方向准确率0.677,位置RSA达0.192,比随机编码提升6.6倍。
- 适合研究具身智能、世界模型与无监督表征学习的学者。
在无语言监督条件下,世界模型从物理探索中学习到什么?我们提出答案由单一原则驱动:物理世界的几何结构。在随机具身探索数据上训练基于VAE的世界模型,发现其潜空间发展出与物理几何一致的空间语义结构——方向准确率为0.677±0.029,显著优于随机初始化编码器的0.547;位置RSA为0.192±0.047,较随机编码器的0.029提升6.6倍,表明训练带来了超越卷积网络先验的真实结构组织。在20个时间检查点上,预测性能与语义对齐同步提升(Spearman r=-0.61, p=0.004),符合共享驱动假说。通过双重消融实验验证:标准KL正则化(beta=0.1)使编码器偏离几何结构,预测性能与语义对齐在第5万步同时跌至近随机水平,完全符合共享驱动预测;将beta降至0.001后,几何可访问性与双能力同时恢复。这些发现确立物理世界几何为世界模型表征的组织原则,对语义扎根的具身智能体设计具有直接启示。
原文摘要 · Abstract (English)
What does a world model learn from physical exploration, without any linguistic supervision? We argue the answer is organized by a single principle: the geometric structure of the physical world. Training a VAE-based world model on random embodied exploration, we find that its latent space develops spatial semantic structure that mirrors physical geometry -- direction accuracy 0.677+-0.029 versus 0.547 for a randomly initialized encoder, and position RSA 0.192+-0.047 versus 0.029 for random encoders (6.6x improvement), showing that training induces genuine structural organization beyond CNN inductive bias. Across 20 temporal checkpoints, prediction performance and semantic alignment co-improve (Spearman r=-0.61, p=0.004), consistent with the shared-driver account. We confirm this through a double knockout: standard KL regularization (beta=0.1) forces the encoder away from geometric structure, and both prediction performance and semantic alignment collapse simultaneously to near-chance by step 50,000 -- exactly as the shared-driver account predicts. Reducing beta to 0.001 restores geometric access and recovers both capabilities together. These findings establish physical world geometry as the organizing principle of world model representations, with direct implications for the design of semantically grounded embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。