让智能体自主进化,通过树状强化学习提升复杂任务表现
SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
- 用树搜索增强策略优化,将稀疏奖励转化为密集中间信号
- 在ALFWorld上达85.07%(文本)和46.27%(多模态)准确率
- 无需真实奖励即可自适应进化,适合长期任务的智能体研究
自我演化是具身智能体在长时序、现实任务中自主改进推理与行为的关键。尽管强化微调(RFT)在提升大模型推理能力上表现优异,但其在具身场景下支持多模态交互的自演化潜力仍待探索。当前面临两大挑战:(i)多步推理任务中缺乏可获取的中间奖励,导致学习信号不足;(ii)依赖人工设计奖励函数,限制了对新任务和环境的泛化能力。为此,我们提出首个面向具身智能体自演化能力的RFT框架——SEEA-R1。为将稀疏延迟奖励转化为更密集的中间信号以改善多步推理,我们提出树状组相对策略优化(Tree-GRPO),将蒙特卡洛树搜索融入GRPO。为实现跨任务与场景的奖励泛化,支持自主适应与奖励驱动的自演化,我们进一步引入多模态生成式奖励模型(MGRM)。在ALFWorld基准上的全面评估显示,SEEA-R1分别取得85.07%(文本)和46.27%(多模态)的性能,超越现有先进方法,包括GPT-4o。在无真实奖励条件下,仍达到80.3%(文本)和44.03%(多模态)的成绩,显著优于所有开源基线,凸显其作为可扩展自演化具身智能体的潜力。额外实验与定性分析进一步验证了该框架在可扩展具身智能研究中的前景。
原文摘要 · Abstract (English)
Self-evolution, the ability of agents to autonomously improve their reasoning and behavior, is essential for the embodied domain with long-horizon, real-world tasks. Despite current advancements in reinforcement fine-tuning (RFT) showing strong performance in enhancing reasoning in LLMs, its potential to enable self-evolving embodied intelligence with multi-modal interactions remains largely unexplored. Specifically, reinforcement fine-tuning faces two fundamental obstacles in embodied settings: (i) the lack of accessible intermediate rewards in multi-step reasoning tasks limits effective learning signals, and (ii) reliance on hand-crafted reward functions restricts generalization to novel tasks and environments. To address these challenges, we present Self-Evolving Embodied Agents-R1, SEEA-R1, the first RFT framework designed for enabling the self-evolving capabilities of embodied agents. Specifically, to convert sparse delayed rewards into denser intermediate signals that improve multi-step reasoning, we propose Tree-based group relative policy optimization (Tree-GRPO) integrates Monte Carlo Tree Search into GRPO. To generalize reward estimation across tasks and scenes, supporting autonomous adaptation and reward-driven self-evolution, we further introduce Multi-modal Generative Reward Model (MGRM). To holistically evaluate the effectiveness of SEEA-R1, we evaluate on the ALFWorld benchmark, surpassing state-of-the-art methods with scores of 85.07% (textual) and 46.27% (multi-modal), outperforming prior models including GPT-4o. SEEA-R1 also achieves scores of 80.3% (textual) and 44.03% (multi-modal) without ground truth reward, surpassing all open-source baselines and highlighting its scalability as a self-evolving embodied agent. Additional experiments and qualitative analysis further support the potential of SEEA-R1 for future research in scalable embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。