arXiv:2605.30542cs.AI2026-05被引 1

让机器人世界模型更可信:用物理结构回答操作问题,而非只预测视觉画面。

Physically Viable World Models: A Case for Query-Conditioned Embodied AI

论文配图:Physically Viable World Models: A Case for Query-Conditioned Embodied AI
图 1 · 摘自论文原文
  • 按干预查询构建物理结构化模型,而非仅预测未来视觉。
  • 在相同外观下,不同物理系统会因结构差异导致行为分歧。
  • 适合需安全决策的机器人规划、控制与验证场景。

具身智能的世界模型必须具备物理可行性:其设计应能通过表征决定行动结果的物理结构来回答干预查询,而非仅预测未来观测。现有观测预测型世界模型虽生成视觉上合理但物理错误的推演。这种失败是结构性的——不同物理系统可能外观相同,但在干预下表现不同。我们通过受控基准测试暴露此问题,固定可见场景而改变潜在物理属性。结果显示,此类模型可能推荐不可行动作、误判交互结果或认证危险行为。我们认为具身智能需要能识别最简物理抽象以回应干预查询的世界模型。该模型由环境表征、隐状态与参数估计、动作定义、干预动力学及查询级响应等模块组成。自主协调器应根据查询选择相关抽象,并组合可兼容的学习与结构化组件。当闭式物理模型不可用、不确定或成本高时,转换模型可为解析式、仿真式、学习式或混合式,但必须保留决定干预结果的结构。该分解使模型可解释、组件可验证、输出可审计,并为新模型设计提供原则,也为旧模型提供可行性检验标准:正确抽象并非最详细的模型,而是保留查询相关区别的最简模型。我们在现有系统无法正确回答的查询上验证了该方法,并概述了协调器如何动态组装和适应物理可行模型用于规划、控制与验证。

原文摘要 · Abstract (English)

World models for embodied AI must be physically viable: constructed to answer intervention queries by representing the physical structure governing action outcomes, rather than merely predicting future observations. Existing observation-predictive world models can produce visually plausible but physically wrong rollouts. This failure is structural; distinct physical systems can look identical yet diverge under intervention. We expose this problem with controlled benchmarks that fix the visible scene while varying latent physics. We show that such models may recommend infeasible actions, mispredict interaction outcomes, or certify unsafe behavior. We argue that embodied AI requires world models that identify the simplest physical abstraction sufficient to answer an intervention query. Such a model comprises modular components, including environment representation, latent state and parameter estimation, action specification, interventional dynamics, and query-level response. An autonomous orchestrator should identify the relevant abstraction and compose compatible learned and structured components per query. When closed-form physics is unavailable, uncertain, or costly, the transition model may be analytic, simulated, learned, or hybrid, but it must preserve the structure that determines interventional outcomes. This decomposition makes the model interpretable, its components verifiable, and its outputs auditable against the query. It also provides a design principle for new world models and a feasibility test for existing ones: the right abstraction is not the most detailed model of the world, but the simplest model that preserves the distinctions relevant to the query. We demonstrate this approach on queries that existing systems fail to answer correctly, and outline how an orchestrator can dynamically assemble and adapt physically viable models for planning, control, and verification.

具身智能世界模型物理可行干预推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。