通过优化网页元素偏好,提升大模型导航准确率
WEPO: Web Element Preference Optimization for LLM-based Web Navigation
- 用非显著元素做负样本,基于距离采样进行无监督偏好学习
- 在Mind2Web上比WebAgent高13.8%,比CogAgent高5.3%
- 适合想提升网页自动化任务效果的研究者和开发者
自主网页导航的快速发展得益于将预训练大语言模型(LLM)作为智能体。然而,现有研究尚未充分利用HTML元素的冗余性进行对比学习。本文提出一种名为网页元素偏好优化(WEPO)的新方法,通过采样基于距离的非显著网页元素作为负样本,结合直接偏好优化(DPO)中的最大似然目标,实现无监督偏好学习。我们在Mind2Web基准上评估了WEPO,实证表明该方法能更有效地将用户高层意图与输出动作对齐。结果表明,本方法达到当前最佳性能,相较WebAgent提升13.8%,相较视觉语言模型CogAgent提升5.3%。研究结果凸显了偏好优化在网页导航及其他基于网页的任务中的潜力,为未来研究提供了有前景的方向。
原文摘要 · Abstract (English)
The rapid advancement of autonomous web navigation has significantly benefited from grounding pretrained Large Language Models (LLMs) as agents. However, current research has yet to fully leverage the redundancy of HTML elements for contrastive training. This paper introduces a novel approach to LLM-based web navigation tasks, called Web Element Preference Optimization (WEPO). WEPO utilizes unsupervised preference learning by sampling distance-based non-salient web elements as negative samples, optimizing maximum likelihood objective within Direct Preference Optimization (DPO). We evaluate WEPO on the Mind2Web benchmark and empirically demonstrate that WEPO aligns user high-level intent with output actions more effectively. The results show that our method achieved the state-of-the-art, with an improvement of 13.8% over WebAgent and 5.3% over the visual language model CogAgent baseline. Our findings underscore the potential of preference optimization to enhance web navigation and other web page based tasks, suggesting a promising direction for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。