让旧轨迹更实用:用查询条件重用长时序智能体经验
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories

- 提出查询条件重用机制,将过往轨迹转化为可复用的靶向支持
- 在多个场景中成功率达62.3%,比完整轨迹高10.7点,少用48.9%令牌
- 适合需要高效复用历史经验的长周期任务系统设计者
检索能找出可能相关的过往轨迹,但无法说明智能体在用户、实体、约束或环境状态变化后如何使用。我们识别出检索后的重用步骤是长时序轨迹记忆的关键瓶颈,并构建了评估框架:固定候选检索、目标状态、模型、解码策略和工具预算,仅改变对智能体的支持方式。我们以查询条件重用(QCR)为例,这是一种简单的目标绑定记录,包含可复用流程、恢复绑定、适用条件与验证要求。QCR用于检验重用假设,而非宣称最优记忆格式。在WebArena、WorkArena和AppWorld共2,391个目标实例上,QCR平均成功率达62.3%,较完整轨迹提升10.7点,且在线令牌消耗减少48.9%。摘要重排序在94.8%的目标中选出可复用记忆,使最终任务成功率与理想选择相差仅1.8点。按轨迹长度和源-目标绑定偏移分析显示,直接注入轨迹随长度增加或源值变更而效用下降,而目标绑定支持则保留更大收益。该框架将检索质量与将经验转化为安全有用支持的问题分离。
原文摘要 · Abstract (English)
Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying the support delivered to the agent. We instantiate the framework with query-conditioned reuse (QCR), a deliberately simple target-bound note that records a reusable procedure, bindings to recover, applicability conditions, and verification requirements. QCR serves to test the reuse hypothesis rather than to claim a universally preferred memory format. Across 2,391 target instances in WebArena, WorkArena, and AppWorld, QCR reaches 62.3% average Success, 10.7 points above Full Trajectory, while using 48.9% fewer online tokens. Summary reranking selects a reusable memory for 94.8% of targets, placing end-task Success within 1.8 points of an oracle reusable selector. Analyses by trajectory length and source--target binding shift show that direct trajectory injection loses much of its utility as traces grow longer or source-specific values change, whereas target-bound support preserves a larger share of the measured gain. The resulting framework separates retrieval quality from the problem of turning retrieved experience into safe, useful support for a new task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。