让视觉语言模型在多轮任务中更聪明地决策。
Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning

- 提出混合优势估计方法,兼顾每步和每轮的优化目标。
- 统一评论家模型可同时估计两种层级的奖励,提升成功率至91%。
- 适合研究多轮智能体强化学习与视觉语言模型决策能力的学者。
大型视觉-语言模型(VLMs)如今可在交互式环境中充当智能体,其成功依赖于跨轮次的连贯推理与决策。尽管端到端训练能提升多轮决策能力,现有方法主要依赖拼接令牌轨迹的逐标记优化,或采用轮内均等信用分配的逐轮优化。本文建立了两类优化的理论框架,推导出同时满足两个目标的混合优势。进一步证明,在恰当选择折扣因子与学习目标下,统一评论家模型可同时估计轮级与标记级价值。据此,提出HyGAE框架,通过混合优势与统一评论家联合优化标记与轮级目标。在五个多轮决策环境上的实验表明,其平均成功率达91%,相比其他方法显著提升10%。深入分析显示,混合优势与回报的精确解析形式对优化至关重要。
原文摘要 · Abstract (English)
Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end training in agentic environments can improve such multi-turn decision-making abilities, current methods mainly rely on either token-wise optimization over concatenated token trajectories or turn-wise optimization with uniform within-turn credit. In this work, we establish theoretical formulations for the two levels of optimization and derive a hybrid advantage that serves both objectives. Furthermore, with an appropriate choice of discount factor and learning target, we prove that a unified critic model can estimate values for both turn-wise and token-wise. As such, we propose HyGAE, an actor-critic framework that jointly optimizes token- and turn-level objectives with the hybrid advantage and unified critic. We conduct extensive evaluations of HyGAE across five multi-turn decision-making environments, where it achieves an average success rate of 91% and a significant improvement of 10% over other methods. Furthermore, we provide an in-depth analysis showing that the exact analytic form of the hybrid advantage and return is crucial for optimization. Project Page: https://wx-zhang.github.io/hygae-web/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。