arXiv:2505.08303cs.CL2025-05被引 1

大模型越强,提示词优化效果越差,黑盒优化在超大规模模型上几乎失效

Evaluating the Effectiveness of Black-Box Prompt Optimization as the Scale of LLMs Continues to Grow

  • 在720亿以上模型上测试三种黑盒提示优化方法
  • 模型越大,优化带来的性能提升越小,出现反向缩放现象
  • 适合研究大模型能力边界或提示工程局限性的读者

黑盒提示优化方法被视为提升大语言模型(LLM)任务表现的有效策略。然而,现有研究多集中于中小型模型(如7B、14B)或早期版本(如GPT-3.5)。随着模型规模持续扩大,例如DeepSeek V3(671B)和Gemini 2.0 Flash,这些优化技术是否仍能带来显著收益尚不明确。本文选取三种代表性黑盒优化方法,在大型LLM上针对四个NLU与NLG数据集进行评估。结果表明,这些方法在超大规模模型上的改进极为有限。进一步分析发现,模型规模是导致优化效果下降的主要因素。通过在Qwen 2.5系列(7B至72B)不同规模模型上实验,观察到提示优化效果随模型增大而减弱,呈现反向缩放规律。

原文摘要 · Abstract (English)

Black-Box prompt optimization methods have emerged as a promising strategy for refining input prompts to better align large language models (LLMs), thereby enhancing their task performance. Although these methods have demonstrated encouraging results, most studies and experiments have primarily focused on smaller-scale models (e.g., 7B, 14B) or earlier versions (e.g., GPT-3.5) of LLMs. As the scale of LLMs continues to increase, such as with DeepSeek V3 (671B), it remains an open question whether these black-box optimization techniques will continue to yield significant performance improvements for models of such scale. In response to this, we select three well-known black-box optimization methods and evaluate them on large-scale LLMs (DeepSeek V3 and Gemini 2.0 Flash) across four NLU and NLG datasets. The results show that these black-box prompt optimization methods offer only limited improvements on these large-scale LLMs. Furthermore, we hypothesize that the scale of the model is the primary factor contributing to the limited benefits observed. To explore this hypothesis, we conducted experiments on LLMs of varying sizes (Qwen 2.5 series, ranging from 7B to 72B) and observed an inverse scaling law, wherein the effectiveness of black-box optimization methods diminished as the model size increased.

提示优化大模型反向缩放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。