让大模型通过不断试错自主构建世界模型,提升决策效率
Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction
- 用动态奖励调整和交互频次递减,让模型主动精简思考过程
- 在迷宫、推箱子等任务中实现单步解题,比以往快数倍
- 适合需要高效推理与迁移能力的智能体研究者
构建稳健的世界模型推理能力对大语言模型智能体在复杂环境中规划与交互至关重要。尽管多轮交互可通过真实反馈提供对环境动态的深入理解,但现有方法常采用僵化的推理流程,限制了模型的主动学习,最终阻碍高效的世界模型推理。为此,我们提出通过高效交互与主动推理(WMAct)实现世界模型内化,使模型摆脱结构化推理束缚,直接通过行动塑造思维。该方法包含两个关键机制:(1) 奖励重缩放机制,根据动作有效性调整结果奖励,激励减少冗余、增强目的性交互;(2) 交互频率退火策略,逐步降低允许的最大交互轮次,迫使模型压缩学习过程、内化环境规律而非过度依赖外部提示。在推箱子、迷宫、出租车等任务上的实验表明,WMAct可实现仅需单轮交互即完成任务的高效世界模型推理,并显著提升在多个推理基准上的表现与复杂环境迁移能力。
原文摘要 · Abstract (English)
Developing robust world model reasoning is crucial for large language model (LLM) agents to plan and interact in complex environments. While multi-turn interaction offers a superior understanding of environmental dynamics via authentic feedback, current approaches often impose a rigid reasoning process, which constrains the model's active learning, ultimately hindering efficient world model reasoning. To address these issues, we explore world-model internalization through efficient interaction and active reasoning (WMAct), which liberates the model from structured reasoning, allowing the model to shape thinking directly through its doing, and achieves effective and efficient world model reasoning with two key mechanisms: (1) a reward rescaling mechanism adjusting outcome reward based on action efficacy to incentivize redundancy reduction and purposeful interaction; (2) an interaction frequency annealing strategy to progressively reduce the maximum allowed interaction turns, which compels the model to condense its learning and internalize environmental dynamics rather than over-relying on environmental cues. Our experiments on Sokoban, Maze, and Taxi show that WMAct yields effective world model reasoning capable of resolving tasks in a single turn that previously required multiple interactions and fosters strong transferability to complex environments, improving performance on a suite of reasoning benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。