优化上下文管理可提升网页导航智能体在未知场景下的泛化能力
From Context to Action: Analysis of the Impact of State Representation and Context on the Generalization of Multi-Turn Web Navigation Agents
- 通过分析交互历史与页面表示对上下文的影响,改进导航策略
- 在未见过的网站、类别和地理位置上性能显著提升
- 适合研究大模型应用泛化与人机交互的开发者与研究人员
基于大语言模型(LLM)的框架已拓展至复杂现实应用,如交互式网页导航。这类系统通过多轮对话响应用户指令,在浏览器中完成任务,带来创新机遇的同时也面临挑战。尽管已有对话式网页导航基准,但对影响智能体性能的关键上下文组件仍缺乏深入理解。本研究旨在填补这一空白,分析网页导航智能体运作中至关重要的上下文要素。我们重点研究上下文管理优化,聚焦交互历史与网页表示的影响。结果表明,通过有效上下文管理,智能体在分布外场景(包括未见过的网站、类别及地理区域)中的表现显著提升。这些发现为LLM驱动智能体的设计与优化提供了重要参考,有助于实现更精准高效的现实世界网页导航。
原文摘要 · Abstract (English)
Recent advancements in Large Language Model (LLM)-based frameworks have extended their capabilities to complex real-world applications, such as interactive web navigation. These systems, driven by user commands, navigate web browsers to complete tasks through multi-turn dialogues, offering both innovative opportunities and significant challenges. Despite the introduction of benchmarks for conversational web navigation, a detailed understanding of the key contextual components that influence the performance of these agents remains elusive. This study aims to fill this gap by analyzing the various contextual elements crucial to the functioning of web navigation agents. We investigate the optimization of context management, focusing on the influence of interaction history and web page representation. Our work highlights improved agent performance across out-of-distribution scenarios, including unseen websites, categories, and geographic locations through effective context management. These findings provide insights into the design and optimization of LLM-based agents, enabling more accurate and effective web navigation in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。