提出轨迹长度感知的对比学习框架,提升GUI智能体的训练效率与稳定性。
Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

- 引入轨迹级质量信号,区分成功与失败路径的精细差异。
- 在相同结果类别内优化执行长度,使成功路径更简洁高效。
- 适合需要高精度自动化交互的GUI任务研究者使用。
由多模态大语言模型驱动的图形用户界面(GUI)智能体在跨数字环境的任务自动化中展现出巨大潜力,强化学习(RL)已成为主流训练范式。然而,广泛使用的组相对策略优化(GRPO)方法存在奖励-梯度错位问题,导致优化效率低下且不稳定。近期工作通过将强化学习重构为可验证奖励(RLVR)的对比或分类目标,提升了稳定性。但现有对比式RLVR方法主要依赖结果级监督,难以捕捉同一结果类别内的轨迹质量细微差异。本文提出针对GUI智能体的长度感知对比学习(LACL-GUI),将轨迹级质量信号融入策略优化。LACL-GUI在成功与失败轨迹中引入结构化偏好,鼓励简洁的成功执行,并根据偏离成功轨迹的程度区分失败质量,同时保持优化稳定性。在GUI智能体基准上的实验表明,LACL-GUI提供了更有效的学习信号,持续优于先前方法,凸显了轨迹级监督在对比式RLVR中的价值。
原文摘要 · Abstract (English)
Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) have shown strong potential for automating tasks across diverse digital environments, where reinforcement learning (RL) has become a dominant training paradigm. However, widely used methods such as Group Relative Policy Optimization (GRPO) suffer from reward-gradient misalignment, leading to inefficient and unstable optimization. Recent work addresses this issue by reformulating RL with verifiable rewards (RLVR) as contrastive or classification-based objectives, which improve stability by eliminating problematic gradient behaviors. Despite this progress, existing contrastive RLVR methods rely primarily on outcome-level supervision and fail to capture fine-grained differences in trajectory quality within the same outcome category. In this paper, we propose Length-Aware Contrastive Learning for GUI Agents (LACL-GUI), a contrastive RLVR framework that incorporates trajectory-level quality signals into policy optimization. LACL-GUI introduces structured preferences within both successful and failed trajectories, encouraging concise successful executions and differentiating failure quality based on divergence from successful trajectories, while preserving optimization stability. Experiments on GUI agent benchmarks show that LACL-GUI provides more effective learning signals and consistently improves agent performance over prior methods, highlighting the value of trajectory-level supervision in contrastive RLVR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。