揭示提示优化中性能与稳定性权衡的新现象,提出可控分析框架MAGE。
MAGE: Understanding Stability-Performance Trade-offs in Multi-component Prompt Optimization
- 构建含记忆、多目标选择和自适应评估的可控分析框架
- 发现优化信号耦合效应:性能提升伴随方差放大,且在多样性增加时更明显
- 适用于研究提示优化组件交互,尤其适合关注系统稳定性的研究者
我们通过MAGE(记忆增强的目标导向提示演化)框架,系统研究迭代提示优化中各组件的相互作用。MAGE并非追求绝对性能最优,而是作为控制性消融平台,整合了情景记忆、多目标帕累托选择和自适应评估。实验发现一种此前未报告的现象——提示优化耦合效应(POCE):当多个随机优化信号在闭环反思中协同工作时,会同时提升性能并放大方差,该行为无法通过孤立分析组件预测。主要发现包括:1)基于失败反馈的反思至关重要,仅依赖评分(OPRO)或抽象批判(Self-Refine)的方法无法改进提示;2)在GSM8K-Hard上,MAGE达46.4%,显著优于GEPA的34.0%(+12.4%,P(MAGE>GEPA)=0.998,5种子,gpt-4o-mini),方差相近(7.3% vs. 7.0%);3)扩大候选池从n=3到n=5,平均准确率提升21.6%,方差增至3.7倍。在Llama 3.1 8B上验证发现,POCE具有头空间依赖性:当基模型已高准确时,方差放大消失。低数据场景(Ntrain=30)下,设计良好的固定提示优于所有反思型优化器,表明支架选择胜过优化器选择。结果表明,提示优化系统应视为耦合随机过程,需兼顾性能与稳定性评估。
原文摘要 · Abstract (English)
How do different components of iterative prompt optimization interact, and what happens when they are combined? We investigate this through MAGE (Memory-Augmented Goal-directed Prompt Evolution), a controlled analysis framework for studying component interaction in prompt optimization. MAGE is not proposed as a superior optimizer in absolute terms; it integrates episodic memory, multi-objective Pareto selection, and adaptive evaluation as a platform for controlled ablation. Our experiments uncover a previously unreported phenomenon, the Prompt Optimization Coupling Effect (POCE): when multiple stochastic optimization signals operate within a closed reflective loop, they interact in ways that simultaneously improve performance and amplify variance, behavior that cannot be predicted by analyzing components in isolation. Three main findings emerge. First, failure-grounded reflection is essential: methods relying only on scores (OPRO) or abstract critique (Self-Refine) fail to improve prompts. Second, MAGE achieves 46.4% versus GEPA's 34.0% on GSM8K-Hard (+12.4%, P(MAGE>GEPA)=0.998, 5 seeds on gpt-4o-mini), with comparable variance (7.3% vs. 7.0%). Third, increasing candidate diversity reveals the clearest POCE signal: expanding the candidate pool from n=3 to n=5 improves mean accuracy by +21.6% while increasing variance by 3.7x. We further validate on Llama 3.1 8B and show POCE is headroom-dependent: when the base model already achieves high accuracy, variance amplification disappears. Finally, in low-data regimes (Ntrain=30), well-designed fixed prompts outperform all reflective optimizers, indicating that scaffold choice dominates optimizer choice. Our results suggest prompt optimization systems behave as coupled stochastic processes and should be evaluated in terms of both performance and stability, not just peak accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。