提出新基准AFTER,评估大模型在重复任务中的技能迁移能力。
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation

- 构建382个企业级任务的基准,覆盖6类角色和22项技能。
- 单一优化轮次提升性能3.7-6.7分,多模型训练技能跨模型准确率达73.1%。
- 发现部分技能通用性强,部分仅适配特定角色工作流。
程序化记忆在提升大模型代理处理重复性工作任务方面日益重要,但其生成可复用技能的能力仍不明确。本文提出AFTER基准,包含382个真实企业任务,覆盖六类职业角色和22种程序化技能,用于评估技能在任务、角色及模型架构间的迁移能力。该基准提供控制性评估场景,涵盖局部改进、跨任务迁移、跨角色迁移和跨模型泛化。实验表明,程序化记忆在工业流程中持续有效:单轮优化使整体性能提升3.7至6.7分;由多模型执行轨迹演化出的技能在跨模型测试中达到73.1%准确率,优于所有单一模型来源。进一步发现,部分技能具有广泛泛化能力,而另一些则局限于特定角色工作流,在迁移中表现下降。这些结果为生产级代理平台中程序化记忆系统的构建、评估与部署提供了实用指导。
原文摘要 · Abstract (English)
Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce reusable skills remains poorly understood. We introduce AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks, roles, and model backbones. The benchmark includes controlled evaluation settings for local improvement, cross-task transfer, cross-role transfer, and cross-model generalization. Experiments show that procedural memory delivers consistent gains in industrial workflows: a single refinement round improves aggregate performance by 3.7-6.7 points, while skills evolved from diverse multi-model execution traces achieve 73.1% cross-model test accuracy, outperforming all single-model trace sources. We further find that some skills generalize broadly across tasks and models, whereas others become specialized to role-specific workflows and lose effectiveness under transfer. These results provide practical guidance for building, evaluating, and deploying procedural memory systems in production agent platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。