用变分子目标强化学习,让视觉语言模型更高效解决复杂决策任务。
Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning
- 将决策问题重构为变分子目标条件强化学习,优化目标为子目标证据下界。
- 在手机与网页控制等任务中,学习效率和性能均优于现有最先进方法。
- 特别适合长时序、稀疏奖励的复杂现实任务,提升智能体自主性。
当前最先进的强化学习方法使视觉语言模型(VLM)代理能够在无监督环境下通过与在线环境的交互进行学习。然而,这些方法在面对具有稀疏奖励和长时程依赖的复杂现实决策任务时,常表现出学习效率低下。本文提出一种新框架——变分子目标条件强化学习(VSC-RL),显著提升了VLM代理在复杂决策任务中的表现。该框架从根本上区别于现有方法,将决策问题重新建模为一个带有新优化目标——子目标证据下界(SGC-ELBO)的变分子目标条件强化学习问题,其包含两个关键组成部分:(a) 最大化子目标条件回报;(b) 最小化与参考目标条件策略之间的差异。我们从理论上和实验上证明,VSC-RL可在不牺牲性能保障的前提下显著提升学习效率。在包括移动设备控制和网页操控在内的多种挑战性基准测试中,VSC-RL始终优于现有最先进方法,展现出更优的学习效率与性能。
原文摘要 · Abstract (English)
State-of-the-art (SOTA) reinforcement learning (RL) methods have enabled vision-language model (VLM) agents to learn from interaction with online environments without human supervision. However, these methods often struggle with learning inefficiencies when applied to complex, real-world decision-making tasks with sparse rewards and long-horizon dependencies. We propose a novel framework, Variational Subgoal-Conditioned Reinforcement Learning (VSC-RL), advancing the VLM agents in resolving challenging decision-making tasks. Fundamentally distinct from existing methods, VSC-RL reformulates the decision-making problem as a variational subgoal-conditioned RL problem with the newly derived optimization objective, Subgoal Evidence Lower BOund (SGC-ELBO), which comprises two key components: (a) maximizing the subgoal-conditioned return, and (b) minimizing the divergence from a reference goal-conditioned policy. We theoretically and empirically demonstrate that the VSC-RL can efficiently improve the learning efficiency without compromising performance guarantees. Across a diverse set of challenging benchmarks, including mobile device and web control tasks, VSC-RL consistently outperforms existing SOTA methods, achieving superior learning efficiency and performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。