让世界模型的隐空间可解释物理意义,提升自主系统可靠性
Four Principles for Physically Interpretable World Models
- 按物理意图组织隐空间,实现可解释性
- 学习与物理规律对齐的不变/等变表示
- 融合多源监督并分块生成,支持验证与扩展
随着自主系统在开放不确定环境中的部署增多,亟需可信赖的世界模型来可靠预测高维观测。现有世界模型的隐变量缺乏与真实物理量和动力学的直接映射,限制了其在规划、控制和安全验证中的实用性与可解释性。本文提出从‘物理启发’转向‘物理可解释’的世界模型,并提炼出四项核心原则:(1) 根据物理意图组织隐空间;(2) 学习与物理世界对齐的不变与等变表示;(3) 将多种监督形式整合进统一训练流程;(4) 分块生成输出以支持可扩展性与可验证性。我们在两个基准上实验验证了每项原则的价值。本工作为实现并利用世界模型的完全物理可解释性开辟了多个新研究方向。
原文摘要 · Abstract (English)
As autonomous systems are increasingly deployed in open and uncertain settings, there is a growing need for trustworthy world models that can reliably predict future high-dimensional observations. The learned latent representations in world models lack direct mapping to meaningful physical quantities and dynamics, limiting their utility and interpretability in downstream planning, control, and safety verification. In this paper, we argue for a fundamental shift from physically informed to physically interpretable world models - and crystallize four principles that leverage symbolic knowledge to achieve these ends: (1) functionally organizing the latent space according to the physical intent, (2) learning aligned invariant and equivariant representations of the physical world, (3) integrating multiple forms and strengths of supervision into a unified training process, and (4) partitioning generative outputs to support scalability and verifiability. We experimentally demonstrate the value of each principle on two benchmarks. This paper opens several intriguing research directions to achieve and capitalize on full physical interpretability in world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。