arXiv:2502.08943cs.CLcs.AI2025-02ACL被引 4

多生成样本提升大模型评估精度,揭示提示难度与错误模式

Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation

  • 构建分层统计模型,融合提示特征与模型随机性
  • 多轮生成使评分方差降低,正确率估计更准确
  • 可量化每个提示的难度,适合评测构建者优化数据

大语言模型在实际应用中表现出色,但现有评估方法常采用确定性生成或单次采样,忽略模型内在随机性,导致采样方差未被捕捉,评估结果不可靠。本文提出一种分层统计模型,综合考虑基准测试特征与大模型随机性。实验证明,使用多个生成结果能显著提升评分估计的准确性并降低方差。同时,可计算每个提示的$\ ext{P}( ext{correct})$——基于正确率的提示级难度分数,提供细粒度分析。我们还构建了数据地图,可视化提示难度与语义分布,有助于发现错误和保障基准质量。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated significant utility in real-world applications, exhibiting impressive capabilities in natural language processing and understanding. Benchmark evaluations are crucial for assessing the capabilities of LLMs as they can provide a comprehensive assessment of their strengths and weaknesses. However, current evaluation methods often overlook the inherent randomness of LLMs by employing deterministic generation strategies or relying on a single random sample, resulting in unaccounted sampling variance and unreliable benchmark score estimates. In this paper, we propose a hierarchical statistical model that provides a more comprehensive representation of the benchmarking process by incorporating both benchmark characteristics and LLM randomness. We show that leveraging multiple generations improves the accuracy of estimating the benchmark score and reduces variance. Multiple generations also allow us to define $\mathbb P\left(\text{correct}\right)$, a prompt-level difficulty score based on correct ratios, providing fine-grained insights into individual prompts. Additionally, we create a data map that visualizes difficulty and semantics of prompts, enabling error detection and quality control in benchmark construction.

大模型评估多生成提示工程统计建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。