让大模型学会物理常识,避免计划纸上谈兵。
Aligning Agentic World Models via Knowledgeable Experience Learning
- 用环境反馈自动构建符号知识库,分过程与目标两类经验
- 在EB-ALFRED和EB-Habitat上表现超越基线,支持跨模型跨环境迁移
- 无需频繁重训,适合需要灵活适应真实物理世界的智能体
当前大型语言模型存在显著的模态断层:虽拥有丰富语义知识,却缺乏对物理世界恒定法则的程序化理解。因此,尽管这些智能体隐式充当世界模型,其生成的规划常出现物理幻觉——逻辑合理但无法执行。现有对齐方法多依赖资源密集型训练或微调,试图将动态环境规则压缩进静态参数,但此类参数封装本质僵化,难以应对物理动态的开放性变化,需持续昂贵再训练。为此,我们提出WorldMind框架,通过合成环境反馈自主构建符号化世界知识库。该框架统一了过程经验(以预测误差强制物理可行性)与目标经验(以成功轨迹引导任务最优性)。在EB-ALFRED和EB-Habitat上的实验表明,WorldMind相较基线取得显著性能提升,并具备出色的跨模型与跨环境迁移能力。
原文摘要 · Abstract (English)
Current Large Language Models (LLMs) exhibit a critical modal disconnect: they possess vast semantic knowledge but lack the procedural grounding to respect the immutable laws of the physical world. Consequently, while these agents implicitly function as world models, their simulations often suffer from physical hallucinations-generating plans that are logically sound but physically unexecutable. Existing alignment strategies predominantly rely on resource-intensive training or fine-tuning, which attempt to compress dynamic environmental rules into static model parameters. However, such parametric encapsulation is inherently rigid, struggling to adapt to the open-ended variability of physical dynamics without continuous, costly retraining. To bridge this gap, we introduce WorldMind, a framework that autonomously constructs a symbolic World Knowledge Repository by synthesizing environmental feedback. Specifically, it unifies Process Experience to enforce physical feasibility via prediction errors and Goal Experience to guide task optimality through successful trajectories. Experiments on EB-ALFRED and EB-Habitat demonstrate that WorldMind achieves superior performance compared to baselines with remarkable cross-model and cross-environment transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。