arXiv:2505.22942cs.CLcs.AI2025-05Conference of the …被引 8

用强化学习提升大模型网页代理的推理能力,让其更懂企业场景下的复杂操作。

WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning

  • 通过规则化奖励机制训练,让模型自主学会每一步的推理与规划。
  • 在WorkArena上比监督微调方法性能提升10.26%至16.59%。
  • 无需人工标注中间步骤,适合企业级自动化任务部署。

基于大语言模型(LLM)的网页代理可在企业环境中自动化执行复杂的实时网页导航任务。然而,依赖监督微调(SFT)的现有代理在面对网页交互固有的动态性时,常因推理能力不足而泛化性与鲁棒性较差。本研究提出WorkForceAgent-R1,一种基于规则式R1风格强化学习框架训练的LLM网页代理,旨在增强面向业务场景的单步推理与规划能力。我们设计了结构化奖励函数,评估输出格式合规性与动作正确性,使WorkForceAgent-R1能隐式学习稳健的中间推理过程,无需显式标注或大量专家示范。在WorkArena基准上的大量实验表明,WorkForceAgent-R1相比SFT基线显著提升10.26%-16.59%,在职场导向的网页导航任务中达到与专有LLM代理(gpt-4o)相当的性能。

原文摘要 · Abstract (English)

Large language models (LLMs)-empowered web agents enables automating complex, real-time web navigation tasks in enterprise environments. However, existing web agents relying on supervised fine-tuning (SFT) often struggle with generalization and robustness due to insufficient reasoning capabilities when handling the inherently dynamic nature of web interactions. In this study, we introduce WorkForceAgent-R1, an LLM-based web agent trained using a rule-based R1-style reinforcement learning framework designed explicitly to enhance single-step reasoning and planning for business-oriented web navigation tasks. We employ a structured reward function that evaluates both adherence to output formats and correctness of actions, enabling WorkForceAgent-R1 to implicitly learn robust intermediate reasoning without explicit annotations or extensive expert demonstrations. Extensive experiments on the WorkArena benchmark demonstrate that WorkForceAgent-R1 substantially outperforms SFT baselines by 10.26-16.59%, achieving competitive performance relative to proprietary LLM-based agents (gpt-4o) in workplace-oriented web navigation tasks.

大模型网页代理强化学习企业自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。