arXiv:2505.22050cs.AIcs.LG2025-05被引 13

用强化学习提升智能体在动态环境中的多步决策能力

Reinforced Reasoning for Embodied Planning

  • 通过监督微调和规则奖励函数,让模型学会结构化决策
  • 在Embrench基准上超越GPT-4o-mini等大模型,泛化能力更强
  • 适合研究具身智能、长时序规划的开发者和研究者

具身规划要求智能体基于动态视觉输入和自然语言目标做出连贯的多步决策。尽管近期视觉语言模型(VLM)在静态感知任务中表现优异,但在交互环境中仍面临时间推理、空间理解与常识推理的挑战。本文提出一种强化微调框架,将R1风格的推理增强引入具身规划。首先从强大闭源模型中蒸馏高质量数据,并进行监督微调(SFT),赋予模型结构化决策先验;随后设计针对多步动作质量的规则奖励函数,采用广义强化偏好优化(GRPO)优化策略。方法在Embrench——一个涵盖域内与域外场景的最新交互式具身任务基准上评估。实验表明,该方法显著优于同等或更大规模的模型,包括GPT-4o-mini及70B以上开源基线,并展现出对未见环境的强大泛化能力。本工作揭示了强化驱动推理在推进具身人工智能长时序规划方面的潜力。

原文摘要 · Abstract (English)

Embodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and natural language goals. While recent vision-language models (VLMs) excel at static perception tasks, they struggle with the temporal reasoning, spatial understanding, and commonsense grounding needed for planning in interactive environments. In this work, we introduce a reinforcement fine-tuning framework that brings R1-style reasoning enhancement into embodied planning. We first distill a high-quality dataset from a powerful closed-source model and perform supervised fine-tuning (SFT) to equip the model with structured decision-making priors. We then design a rule-based reward function tailored to multi-step action quality and optimize the policy via Generalized Reinforced Preference Optimization (GRPO). Our approach is evaluated on Embench, a recent benchmark for interactive embodied tasks, covering both in-domain and out-of-domain scenarios. Experimental results show that our method significantly outperforms models of similar or larger scale, including GPT-4o-mini and 70B+ open-source baselines, and exhibits strong generalization to unseen environments. This work highlights the potential of reinforcement-driven reasoning to advance long-horizon planning in embodied AI.

具身智能多步决策强化学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。