MELLON提升网页导航任务准确率9.26%,强化图文协同推理能力。
MELLON - Multimodal Enhanced LLM for Online Navigation

- 融合文本与图像的多模态增强框架,提升导航决策能力
- 单轮训练后任务完成率提升9.26%,在WebShop上表现优异
- 适合需要强多模态理解的自动化网页交互研究者
网页导航智能体可处理不同网站上的多种任务。当前基线方法或为单模态,或在面对多模态输入时缺乏强推理能力。针对真实场景网站模拟基准WebShop,我们探索了文本与图像的对齐,以及多模态推理与规划能力,以提升网页导航智能体性能。提出三种创新多模态增强方案:MELLON、VQAgent和多模态排序器。MELLON在仅经过一轮训练后,任务完成准确率提升9.26%。研究结果表明,需进一步探索多模态方法,尤其应加强训练规模与对齐策略,以提高网页导航智能体的有效性。
原文摘要 · Abstract (English)
Web navigation agents are capable of addressing various types of tasks on different websites. Current baselines on web navigation are either unimodal or lack strong reasoning abilities given multimodal inputs. Focusing on the WebShop benchmark, a real-world website simulation, we explore the alignment of text and images, as well as multimodal reasoning and planning abilities, to enhance the performance of web navigation agents. We propose three innovative multimodal enhancements: Multimodal Enhanced LLM for Online Navigation (MELLON), VQAgent, and Multimodal Ranker. MELLON demonstrates a significant improvement in task completion accuracy, with a 9.26% increase after just one epoch of training. Our findings suggest the necessity of further exploration into multimodal approaches, with a focus on more extensive training and alignment strategies to enhance the effectiveness of web navigation agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。