在线网页代理的技能与记忆模块未必值回票价,预算匹配下基础模型更优。
Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents
- 在固定推理预算下,对比增强模块与等预算基础模型
- 四类任务中,基础模型成功率更高或相当,且总耗token更少
- 运行波动影响显著,应作为评估核心指标
在线网页代理常通过记忆、工作流或技能模块增强基础智能体。这些模块虽能提升性能,但会增加测试时的token消耗,该成本常被忽视。本文研究在线增强模式,即每项任务都需支付额外开销,并在固定总推理预算下重新评估其价值。我们对比了AWM、ASI和ReasoningBank三种方法,与使用相同预算进行额外智能体步骤的基线模型。在四个WebArena领域及三个模型(Gemini 3 Flash、GPT-5.4-mini、Qwen 3.6-27B)上,基线模型在整体成功率达标或超越所有增强方法,且常以更少的总token完成。在WorkArena-L1上,使用Qwen 3.6-27B也呈现类似趋势,表明该结论扩展至企业级知识工作场景。结果表明,技能与工作流记忆仅在特定领域有效,其表面收益在预算匹配下往往消失。此外,我们还发现运行间方差显著影响结果,应作为在线网页代理评估的核心标准。
原文摘要 · Abstract (English)
Online web agents often augment a base actor with memory, workflow, or skill modules. These modules can improve performance, but they also consume test-time tokens, a cost rarely reported alongside the actor's inference cost. We study online augmentation, where this overhead is paid on every task, and re-evaluate its benefits under a fixed total inference budget. We compare AWM, ASI, and ReasoningBank with a token-matched vanilla baseline that uses the same budget for additional actor steps. Across four WebArena domains and three models, Gemini 3 Flash, GPT-5.4-mini, and Qwen 3.6-27B, the vanilla baseline matches or surpasses all three augmentation methods in aggregate success rate while often using fewer total tokens. We observe a similar trend on WorkArena-L1 with Qwen 3.6-27B, indicating that the effect extends to enterprise knowledge-work tasks. Our results suggest that skills and workflow memory can be useful in specific domains, but their apparent gains often vanish against a budget-matched actor. We further show that run-to-run variance materially affects outcomes and should be reported as a core evaluation criterion for online web agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。