研究提示词排名稳定性,发现选最优提示易受随机因素干扰。
On the Stability of Prompt Ranking in Large Language Model Evaluation
- 用置信下界策略评估提示词性能与方差,提升选择鲁棒性
- 在三个模型、两个任务上,顶级提示常因随机变化而轮换
- 适合做模型评测或提示工程的研究者参考
基于提示词的交互已成为大语言模型使用的主要范式,通过评估多个候选提示并选取排名靠前者用于下游任务。该流程隐含假设:提示词排名在评价条件微小变化下保持稳定。本文系统研究了常见变异源(如随机种子和有限评估子集)对提示词排名稳定性的影响。在三个开源大模型和两个基准任务上,我们发现尽管整体排名相关性通常中等到高,但最佳提示的归属频繁变动,导致选择结果不可靠。为此,我们提出一种基于置信下界的简单稳定性感知选择策略,同时考虑性能与方差。实验表明,该方法在不稳定环境下显著提升鲁棒性,同时在较稳定场景中仍具竞争力。结果强调了在提示词选择与模型评测中必须考虑评估不确定性。
原文摘要 · Abstract (English)
Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evaluated and the top-ranked one is selected for downstream use. This workflow implicitly assumes that prompt rankings are stable under minor variations in evaluation conditions. In this paper, we systematically study prompt ranking stability under common sources of variability, including random seeds and limited evaluation subsets. Across three open-weight LLMs and two benchmark tasks, we find that while overall rank correlations are often moderate to high, the identity of the top-performing prompt frequently changes, leading to unreliable selection decisions. To address this issue, we propose a simple stability-aware selection strategy based on a lower confidence bound, which accounts for both performance and variance. Our results show that this approach improves robustness in unstable settings while remaining competitive in more stable regimes. These findings highlight the importance of accounting for evaluation uncertainty in prompt selection and LLM benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。