用少量人类操作示范,让多模态网页代理快速适应新网站。
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations

- 通过人类示范实现少样本自适应,无需大规模微调。
- 在两个基准上任务成功率提升21%至65%,最高增7.21个百分点。
- 适合需要快速适配私有或专有平台的企业级应用。
当前基于多模态大模型(MLLM)的网页代理虽能自主执行多种任务,但在未见过的网站和领域上表现不佳,限制了其在企业专有平台的应用。本文提出AdaptAgent框架,利用最多2个用户操作示范,使具有私有权重或开源权重的多模态网页代理快速适应新网站。在Mind2Web与VisualWebArena两个基准上的实验表明,使用上下文示范(对私有模型)或元适应示范(对元学习的开源模型),任务成功率较非适配的最先进模型提升3.36%至7.21%,相对增长21.03%至65.75%。分析显示,多模态示范优于纯文本示范,元学习中的数据选择策略影响代理泛化能力,且示范数量越多,成功率越高。结果表明,少样本适应是超越大规模预训练与微调的新发展方向。
原文摘要 · Abstract (English)
State-of-the-art multimodal web agents, powered by Multimodal Large Language Models (MLLMs), can autonomously execute many web tasks by processing user instructions and interacting with graphical user interfaces (GUIs). Current strategies for building web agents rely on (i) the generalizability of underlying MLLMs and their steerability via prompting, and (ii) large-scale fine-tuning of MLLMs on web-related tasks. However, web agents still struggle to automate tasks on unseen websites and domains, limiting their applicability to enterprise-specific and proprietary platforms. Beyond generalization from large-scale pre-training and fine-tuning, we propose building agents for few-shot adaptability using human demonstrations. We introduce the AdaptAgent framework that enables both proprietary and open-weights multimodal web agents to adapt to new websites and domains using few human demonstrations (up to 2). Our experiments on two popular benchmarks -- Mind2Web & VisualWebArena -- show that using in-context demonstrations (for proprietary models) or meta-adaptation demonstrations (for meta-learned open-weights models) boosts task success rate by 3.36% to 7.21% over non-adapted state-of-the-art models, corresponding to a relative increase of 21.03% to 65.75%. Furthermore, our additional analyses (a) show the effectiveness of multimodal demonstrations over text-only ones, (b) shed light on the influence of different data selection strategies during meta-learning on the generalization of the agent, and (c) demonstrate the effect of number of few-shot examples on the web agent's success rate. Overall, our results unlock a complementary axis for developing widely applicable multimodal web agents beyond large-scale pre-training and fine-tuning, emphasizing few-shot adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。