让网页智能体自我发现短板并主动学习,提升复杂环境适应力。
Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration

- 设计三角色对抗机制,自动识别智能体能力边界。
- 在19个真实网站上构建20k条数据集,验证性能显著提升。
- 适合研究自主智能体、多模态模型应用的开发者参考。
多模态大模型(MLLMs)的发展推动了网页智能体的进步。然而,现有智能体常依赖手工设计的执行流程或昂贵的专家轨迹,难以适应复杂动态环境。为此,我们提出SCALE(自认知感知学习与探索)框架,通过选择器、预测器和评判者三个对抗角色,使智能体能自主发现自身局限并通过环境探索拓展认知边界。此外,我们提出SCALE-Hop图探索策略,支持全局规划,避免局部探索陷阱。为支持训练,我们构建了包含19个真实网站、20,000条多样化任务和结构化示范的SCALE-20k数据集,示范来自智能体自身的探索轨迹。实验表明,该方法显著提升了多种MLLM在不同网页环境下的表现与泛化能力。本框架为构建真正自主、可适应的网页智能体提供了可扩展、通用的解决方案。
原文摘要 · Abstract (English)
Recent advances in Multimodal Large Language Models (MLLMs) have led to promising progress in web agents. However, existing web agents often rely on handcrafted execution pipelines or expensive expert trajectories, limiting their adaptability to complex, dynamic environments. To address these challenges, we propose SCALE (Self-Cognitive-Aware Learning and Exploration), which leverages three adversarial roles, Selector, Predictor, and Judger to autonomously discover the agent's limitations and expand its cognitive boundaries through environmental exploration. Moreover, we propose SCALE-Hop, a graph exploration strategy that facilitates global planning and helps agents avoid local exploration traps. To further support learning, we construct SCALE-20k, a large-scale dataset collected from 19 real-world websites, containing diverse task types and structured demonstrations generated from SCALE's exploration traces. Experimental results show that our approach significantly improves the performance and generalization of multiple MLLMs in various web environments. Our framework offers a scalable and generalizable solution for building truly autonomous and adaptive web agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。