用强化学习替代监督微调,仅用5200样本超越百万级数据训练效果
GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- 基于强化学习设计新方法,分解并优化规则型奖励机制
- 提出对抗性KL因子稳定训练,避免奖励过拟合
- 仅需5200样本即超越超百万样本的监督训练结果
图形用户界面视觉定位(GUI-VG)作为GUI智能体的核心能力,传统上依赖多模态大模型的监督微调(SFT),需大量数据标注与高昂训练成本。然而随着大模型在预训练阶段已覆盖GUI领域,后训练阶段的全面SFT必要性逐渐减弱。近期基于规则的强化微调(RFT)展现出高效潜力,但其在GUI-VG中的最优应用方式尚不明确。为此,本文提出GuirlVG,通过系统性实证研究与新型稳定性技术构建强化学习框架。研究发现直接应用RFT性能反而低于SFT基线,因此深入剖析RFT各组件并优化配置:首先分解RFT核心模块,分析其最优形式;其次提出动态稳定机制——对抗性KL因子,抑制奖励过优化;最后探索不同训练配置以提升效果。大量实验表明,GuirlVG仅使用5.2K训练样本,即在ScreenSpot上实现7.7%提升、ScreenSpotPro上17.2%提升,并在ScreenSpotV2上达到91.9%准确率,显著超越需超1000万样本训练的SFT方法。
原文摘要 · Abstract (English)
Graphical user interface visual grounding (GUI-VG), a core capability for GUI agents, has primarily relied on supervised fine-tuning (SFT) of multimodal large language models (MLLMs), which demands extensive data curation and significant training costs. However, as MLLMs continue to advance and even cover GUI domains during pretraining, the necessity of exhaustive SFT post-training becomes increasingly questionable. Meanwhile, recent successes of rule-based reinforcement fine-tuning (RFT) suggest a more efficient alternative. Despite this promise, the optimal manner of applying RFT for GUI-VG remains unexplored. To bridge this gap, we introduce GuirlVG, a reinforcement learning-based GUI-VG method built on a systematic empirical study and a novel stabilization technique. We find that naive application of RFT underperforms the SFT baseline, motivating a deeper exploration. First, we decompose RFT into its core components and analyze the optimal formulation of each. Second, we propose a novel Adversarial KL Factor that dynamically stabilizes training to mitigate reward over-optimization. Third, we further explore the training configurations of RFT to enhance effectiveness. Extensive experiments show that GuirlVG, with only 5.2K training samples, outperforms SFT methods trained on over 10M samples, achieving a 7.7% improvement on ScreenSpot, a 17.2% improvement on ScreenSpotPro, and 91.9% accuracy on ScreenSpotV2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。