用智能筛选替代盲目评估,让自动提示优化更高效准确。
Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees
- 基于题目区分度设计动态评估调度策略,兼顾精准与效率。
- 在相同预算下提升平均准确率6.2%,评估量减少35%-60%。
- 适合追求高效提示优化的科研人员与工程实践者。
自动提示优化(APO)依赖高质量的评估信号,但对每个候选提示在全训练集上评分代价高昂。现有方法或固定评估子集(合理但不敏感提示)、或启发式动态调整(灵活但不稳定且无理论保障)。本文提出提示感知在线评估调度(POES),将提示优化视为在线测试问题:提示为考生,训练样本为试题,调度器应选择最能区分强候选者的题目。POES整合基于项目反应理论的区分度效用、设施位置覆盖项及考虑切换成本的热启动替换机制,目标函数具有单调亚模性,保证冷启动下(1-1/e)近似率,热启动更新漂移有界。自适应控制器根据优化进度调节探索与利用平衡。在三个基准家族共36个任务上,POES在相同评估预算下实现最高平均准确率(较最优基线提升6.2%),仅增加约4%令牌开销。在k=20时的精心筛选表现优于朴素评估k=30-50,令牌消耗降低35%-60%。结果表明,评估调度是APO的核心组件,而非附属细节。
原文摘要 · Abstract (English)
Automatic prompt optimization (APO) hinges on the quality of its evaluation signal, yet scoring every prompt candidate on the full training set is prohibitively expensive. Existing methods either fix a single evaluation subset before optimization begins (principled but prompt-agnostic) or adapt it heuristically during optimization (flexible but unstable and lacking formal guarantees). We observe that APO naturally maps to an online adaptive testing problem: prompts are examinees, training examples are test items, and the scheduler should select items that best discriminate among the strongest candidates. This insight motivates Prompt-Aware Online Evaluation Scheduling (POES), which integrates an IRT-based discrimination utility, a facility-location coverage term, and switching-cost-aware warm-start swaps into a unified objective that is provably monotone submodular, yielding a (1-1/e) greedy guarantee for cold starts and bounded drift for warm-start updates. An adaptive controller modulates the exploration-exploitation balance based on optimization progress. Across 36 tasks spanning three benchmark families, POES achieves the highest overall average accuracy (6.2 percent improvement over the best baseline) with negligible token overhead (approximately 4 percent) at the same evaluation budget. Moreover, principled selection at k = 20 examples matches or exceeds the performance of naive evaluation at k = 30-50, reducing token consumption by 35-60 percent, showing that selecting smarter is more effective than selecting more. Our results demonstrate that evaluation scheduling is a first-class component of APO, not an implementation detail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。