通过重构思维链提升网页智能体的反思、分支与回溯能力。
WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollback
- 将推理过程转化为思维链,显式训练反思、分支与回溯技能。
- 在OpenWebVoyager等基准上显著提升智能体表现,最高提升32%。
- 适合想提升网页代理推理能力的研究者和开发者。
由大语言模型驱动的网页智能体在下一代人工智能中展现出巨大潜力,但其在不确定、动态的网络环境中推理能力有限,制约了可靠部署。本文识别出有效网页智能体所需的关键推理技能:反思与前瞻、分支探索、回滚修正,并通过重建智能体在推理时的决策逻辑,生成链式思维(Chain-of-Thought)示例数据。我们在自改进基准OpenWebVoyager上进行实验,证明仅通过简单微调,将关键推理模式注入主干LLM即可显著提升性能。该方法在WebVoyager、Mind2web-live和SimpleQA(网页搜索)等多个基准上均取得显著进步,验证了针对性强化推理技能对网页智能体的可行性与有效性。
原文摘要 · Abstract (English)
Web agents powered by Large Language Models (LLMs) show promise for next-generation AI, but their limited reasoning in uncertain, dynamic web environments hinders robust deployment. In this paper, we identify key reasoning skills essential for effective web agents, i.e., reflection & lookahead, branching, and rollback, and curate trajectory data that exemplifies these abilities by reconstructing the agent's (inference-time) reasoning algorithms into chain-of-thought rationales. We conduct experiments in the agent self-improving benchmark, OpenWebVoyager, and demonstrate that distilling salient reasoning patterns into the backbone LLM via simple fine-tuning can substantially enhance its performance. Our approach yields significant improvements across multiple benchmarks, including WebVoyager, Mind2web-live, and SimpleQA (web search), highlighting the potential of targeted reasoning skill enhancement for web agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。