优化提示词常如抛硬币,仅在特定任务中有效
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems

- 通过大规模实验验证提示词是否需联合优化
- 只有具备可挖掘输出结构的任务才受益,最高提升6.8分
- 提供快速诊断方法,避免无效优化投入
在复合AI系统中,提示词优化的效果与抛硬币无异:在Claude Haiku 4.5上进行72次优化(6种方法×4个任务×3次重复),49%的结果低于零样本;在Amazon Nova Lite上失败率更高。然而,在一个任务中,所有六种方法均优于零样本,提升最高达+6.8分。我们通过18,000次网格评估和144次优化运行,检验了端到端优化工具(如TextGrad、DSPy)的两个前提:(A) 代理提示词存在交互,需联合优化;(B) 单个提示词值得优化。结果表明,交互效应均不显著(p > 0.52,所有F < 1.0),且优化仅在任务具有可利用输出结构时有效——即模型能生成但默认不输出的格式。进一步机制分析发现,指令微调会将输入表述压缩至狭窄输出分布,消除了联合优化所依赖的表述敏感性。为此,我们提出两阶段诊断:80美元的ANOVA预测试验判断代理耦合,10分钟头寸测试预测优化价值,将随机决策变为有依据的选择。
原文摘要 · Abstract (English)
Prompt optimization in compound AI systems is statistically indistinguishable from a coin flip: across 72 optimization runs on Claude Haiku 4.5 (6 methods $\times$ 4 tasks $\times$ 3 repeats), 49% score below zero-shot; on Amazon Nova Lite, the failure rate is even higher. Yet on one task, all six methods improve over zero-shot by up to $+6.8$ points. What distinguishes success from failure? We investigate with 18,000 grid evaluations and 144 optimization runs, testing two assumptions behind end-to-end optimization tools like TextGrad and DSPy, in the order they must be answered: (A) agent prompts interact, requiring joint rather than independent optimization, and (B) individual prompts are worth optimizing at all. Interaction effects are never significant ($p > 0.52$, all $F < 1.0$), and optimization helps only when the task has exploitable output structure: a format the model can produce but does not default to. We further give a mechanistic account: instruction-tuning compresses input phrasing into a narrow output distribution, eliminating the very phrasing-sensitivity that joint optimization assumes. We provide a two-stage diagnostic: an \$80 ANOVA pre-test for agent coupling, and a 10-minute headroom test that predicts whether optimization is worthwhile, turning a coin flip into an informed decision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。