用随机采样提升大模型评估可靠性,避免因提示词微调导致结果偏差。
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
- 通过扰动提示词的统计方法,评估大模型在不同输入下的稳定性。
- 发现GPT-4o和Claude-3.7-Sonnet等顶级模型仍存在显著提示敏感性。
- 适用于任何模型、任务和指标,提供可复现的评估方案。
大语言模型对提示词措辞高度敏感,但现有基准通常仅使用单一提示词报告性能,引发评估可靠性担忧。本文提出基于语义保持提示扰动空间的随机矩评估方法,给出可靠评估的形式化定义,并引入ReliableEval,用于估计获得有意义结果所需的提示重采样次数。利用该框架,我们对五款前沿大模型进行随机评估,发现即使表现最优的GPT-4o与Claude-3.7-Sonnet也表现出显著的提示敏感性。该方法具备模型、任务和度量无关性,为实现有意义且稳健的大模型评估提供可操作的范式。
原文摘要 · Abstract (English)
LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations. In this work, we argue for a stochastic method of moments evaluation over the space of meaning-preserving prompt perturbations. We introduce a formal definition of reliable evaluation that accounts for prompt sensitivity, and suggest ReliableEval - a method for estimating the number of prompt resamplings needed to obtain meaningful results. Using our framework, we stochastically evaluate five frontier LLMs and find that even top-performing models like GPT-4o and Claude-3.7-Sonnet exhibit substantial prompt sensitivity. Our approach is model-, task-, and metric-agnostic, offering a recipe for meaningful and robust LLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。