arXiv:2504.16073cs.CL2025-04被引 7

用过程奖励引导VLM在界面上实时决策,提升操作准确率和任务成功率。

Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation

  • 推理时引入奖励模型指导VLM每步动作选择,实现动态优化。
  • 静态环境动作准确率提升3.4%,动态环境任务成功率提高约33%。
  • 适合需要高可靠性的GUI自动化任务,如智能助手或测试工具。

视觉语言模型(VLM)在复杂图形用户界面(GUI)交互任务中能力显著提升。然而,现有框架在挑战性界面环境中仍难以生成正确动作。主流商业VLM为黑盒,开源VLM微调需大量资源。此外,现有基于轨迹的评估与优化方法常因反馈延迟和局部最优而失效。为此,我们提出一种在推理阶段通过奖励模型对VLM代理进行过程监督的方法,使其在GUI导航与控制中实时优化每一步动作,从而提升静态与动态环境下的表现。实验显示,在三个GUI导航任务中,该方法使静态环境单步动作准确率提升3.4%,动态环境中任务成功率提升约33%。结合轨迹反思与重试机制后,任务成功率进一步显著提升。

原文摘要 · Abstract (English)

Recent advancements in visual language models (VLMs) have notably enhanced their capabilities in handling complex Graphical User Interface (GUI) interaction tasks. Despite these improvements, current frameworks often struggle to generate correct actions in challenging GUI environments. State-of-the-art commercial VLMs are black-boxes, and fine-tuning open-source VLMs for GUI tasks requires significant resources. Additionally, existing trajectory-level evaluation and refinement techniques frequently fall short due to delayed feedback and local optimization issues. To address these challenges, we propose an approach that guides VLM agents with process supervision by a reward model during GUI navigation and control at inference time. This guidance allows the VLM agent to optimize actions at each inference step, thereby improving performance in both static and dynamic environments. In particular, our method demonstrates significant performance gains in three GUI navigation tasks, achieving a 3.4% improvement in single step action accuracy for static environments, along with a around 33% increase in task success rate in one dynamic environment. With further integration of trajectory reflection and retry mechanisms, we also demonstrate even greater enhancement in task success.

GUI导航VLM强化学习推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。