提出新基准RECAP,测试提示词在持续适应中的表现。
RECAP: Regression Evaluation for Continual Adaptation of Prompts
- 采用先适应后测试的主动协议,仅给约束信息不给测试数据
- 六种方法在五模型三调度下性能无显著提升,延迟更高
- 揭示现有方法不适用于生产中持续演化的合规需求
生产环境中的智能体系统需在每次交互后立即适应不断变化的约束,如工具调用更新合规阈值或政策添加披露要求,容错率极低。当前基准多假设静态约束或反应式评估反馈,缺乏对主动适应场景的评测。本文提出RECAP基准,以严格主动适应-测试协议衡量提示词优化方法在约束层面的持续学习现象(遗忘、退化、正向迁移)。在五个大语言模型和三种约束演化策略下评估六种方法,发现其性能未显著提升,且延迟更高。这些为离线或反应式场景设计的方法,无法应对主动适应需求。研究强调需开发能在部署中应对动态需求的主动提示适应方法。
原文摘要 · Abstract (English)
Production agentic systems routinely face evolving constraints and must comply from the very next interaction. Scenarios like a tool-call notification changing a compliance threshold or a policy update adding disclosure requirements fit this criteria, having close to no room for errors in production. This proactive adaptation setting is common in deployment, but absent from current benchmarks, which assume either static constraint sets or reactive protocols with evaluation feedback. We introduce RECAP, a benchmark that measures continual-learning phenomena (forgetting, regression, forward transfer) at the constraint level under a strictly proactive adapt-then-test protocol: prompt optimization methods receive only the constraint specification and must generalize before seeing any test data. Evaluating six methods across five LLMs and three schedules with evolving constraints, we find that these methods show no significant improvement in performance, even after incurring a higher latency. These methods, designed for offline or reactive settings, are inadequate for the proactive paradigm. Our work emphasizes the growing need for designing proactive prompt adaptation methods, where the models must remain robust to evolving needs in deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。