在规则误导环境下,通过失败反推动态规律,构建可执行世界模型。
Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models

- 用失败更新作为结构信号,发现程序混淆的动态规律
- 在Baba Is You变体上显著提升可执行模型学习效果
- 适合研究自监督学习与具身智能的科研人员
可执行世界模型需准确捕捉环境转移规律而非表面语义捷径。本文研究在先验错位下的在线可执行世界模型学习:智能体仅通过交互数据,无法获取规则说明、奖励信号或可信词汇先验。提出闭环系统Alice,将候选更新失败视为结构性信号:当一个候选解释新转换却丢失旧解释时,保留冲突揭示了当前程序混淆的动态。Alice将这些冲突转化为具有类别区分性的假设类,既提供紧凑的反例以指导更新,又引导探索新且未充分覆盖的转换。在替换规则属性标签为无关词汇的Baba in Wonderland变体上评估,实验表明Alice显著提升先验错位下的可执行模型学习性能,消融实验显示类别精炼与类别感知探索均贡献显著。
原文摘要 · Abstract (English)
Executable world models can be read, edited, executed, and reused for planning, but only if the program captures the environment's transition law rather than semantic shortcuts in its surface vocabulary. We study online executable world-model learning under prior misalignment, where an agent must induce state-dependent dynamics from interaction evidence alone, without rule descriptions, reward signals, or trustworthy lexical priors. We introduce Alice, a closed-loop system that treats failed candidate updates as structural signal: when a candidate explains a new transition but loses previously explained ones, the preservation conflict reveals dynamics that the current program had conflated. Alice refines these conflicts into hypothesis classes that both provide compact, class-stratified preservation counterexamples for update and guide frontier exploration toward transitions that are novel and underrepresented with respect to the current program. We evaluate Alice on Baba in Wonderland, a prior-misaligned variant of Baba Is You that preserves simulator dynamics while replacing semantically meaningful rule-property labels with unrelated words. Experiments show that Alice substantially improves executable world-model learning under prior misalignment, and ablations show that both class refinement and class-aware exploration contribute.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。