用行为一致性提升文本世界模型的长期预测能力
Beyond State Consistency: Behavior Consistency in Text-Based World Models

- 设计行为一致性奖励(BehR),衡量模型预测状态与真实状态下的动作概率差异
- 在WebShop上实现显著长期对齐提升,且单步预测质量不下降
- 适合用于离线评估和推理时规划的场景,尤其在高精度需求下
世界模型在交互式智能体的在线规划与离线评估中日益关键。在文本环境中,传统训练与评估多依赖单步指标如精确匹配,旨在提升预测状态与真实状态的相似性,但此类指标难以捕捉实际代理行为。为此,本文提出一种新的行为对齐训练范式,聚焦于优化一个可计算的步骤级指标——行为一致性奖励(BehR),该指标衡量在固定参考智能体下,真实状态与模型预测状态中下一个动作的似然变化程度。在WebShop与TextWorld上的实验表明,基于BehR的训练在多种设置下提升了长期对齐效果,其中在WebShop上提升最为明显,而在接近性能上限的场景中改善较小;同时在四个设置中的三个里,保持或提升了单步预测质量。使用BehR训练的世界模型在离线替代评估中表现出更低的误报率,并在推理时前瞻规划中获得适度但令人鼓舞的提升。
原文摘要 · Abstract (English)
World models have been emerging as critical components for assessing the consequences of actions generated by interactive agents in online planning and offline evaluation. In text-based environments, world models are typically evaluated and trained with single-step metrics such as Exact Match, aiming to improve the similarity between predicted and real-world states, but such metrics have been shown to be insufficient for capturing actual agent behavior. To address this issue, we introduce a new behavior-aligned training paradigm aimed at improving the functional consistency between the world model and the real environment. This paradigm focuses on optimizing a tractable step-level metric named Behavior Consistency Reward (BehR), which measures how much the likelihood of a logged next action changes between the real state and the world-model-predicted state under a frozen Reference Agent. Experiments on WebShop and TextWorld show that BehR-based training improves long-term alignment in several settings, with the clearest gains in WebShop and less movement in near-ceiling regimes, while preserving or improving single-step prediction quality in three of four settings. World models trained with BehR also achieve lower false positives in offline surrogate evaluation and show modest but encouraging gains in inference-time lookahead planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。