用网页教程增强GUI代理,让AI更懂复杂操作
Retrieval-augmented GUI Agents with Generative Guidelines
- 推理时调用网络教程,补足训练数据不足的短板
- 在三个任务中表现优于基线,提升幅度2.6%至13.3%
- 无需修改模型,可直接接入任意视觉语言代理
基于视觉-语言模型(VLM)的GUI代理在自动化复杂数字任务方面展现出潜力。然而,其在真实场景中的效果常受限于训练数据稀缺及任务本身的复杂性,尤其在涉及罕见、未见场景的长尾知识时表现不佳。本文提出RAG-GUI,一种轻量级VLM,在推理时利用网络教程进行增强。RAG-GUI首先通过监督微调(SFT)进行热启动,再通过自指导拒绝采样微调(RSF)进一步优化。该方法具有模型无关性,可作为通用插件提升任意VLM-based代理性能。在三个不同任务上的评估表明,RAG-GUI持续优于基线代理,在两种模型规模下均比其他推理基线高出2.6%至13.3%,展现出强泛化能力与实际应用场景中的即插即用特性。
原文摘要 · Abstract (English)
GUI agents powered by vision-language models (VLMs) show promise in automating complex digital tasks. However, their effectiveness in real-world applications is often limited by scarce training data and the inherent complexity of these tasks, which frequently require long-tailed knowledge covering rare, unseen scenarios. We propose RAG-GUI , a lightweight VLM that leverages web tutorials at inference time. RAG-GUI is first warm-started via supervised finetuning (SFT) and further refined through self-guided rejection sampling finetuning (RSF). Designed to be model-agnostic, RAG-GUI functions as a generic plug-in that enhances any VLM-based agent. Evaluated across three distinct tasks, it consistently outperforms baseline agents and surpasses other inference baselines by 2.6% to 13.3% across two model sizes, demonstrating strong generalization and practical plug-and-play capabilities in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。