动态调整测试题,让大模型评估更准更快。
Fluid Language Model Benchmarking
- 根据模型水平智能选题,类似自适应考试。
- 用50分之一的题目达成更高准确度和更低波动。
- 适合追求高效精准评估的研究者与工程师。
语言模型评估面临多重挑战:全面评估成本高,基准测试常无法衡量真实能力,且标注错误和基准饱和会降低评估质量。现有方法多孤立解决单一问题,忽视整体评估质量。本文提出流式评估(Fluid Benchmarking),借鉴心理测量学思想,认为测试题的价值取决于模型能力水平,因此评估应随模型动态调整。方法上,基于已有评估结果估计项目反应模型,并据此动态选择测试项,类似教育领域的计算机自适应测试。实验对比随机抽样及基于项目反应理论的基线方法,在效率、有效性、方差和饱和度四个维度均表现更优:例如在MMLU上仅用50倍少的题目即实现更高有效性和更低方差。分析表明,项目反应理论将性能映射到潜在能力空间以提升有效性,动态选题则降低方差。结果表明,摆脱静态评估可显著提升语言模型评测质量。
原文摘要 · Abstract (English)
Language model (LM) benchmarking faces several challenges: comprehensive evaluations are costly, benchmarks often fail to measure the intended capabilities, and evaluation quality can degrade due to labeling errors and benchmark saturation. Although various strategies have been proposed to mitigate these issues, they tend to address individual aspects in isolation, neglecting broader questions about overall evaluation quality. Here, we introduce Fluid Benchmarking, a new evaluation approach that advances LM benchmarking across multiple dimensions. Inspired by psychometrics, Fluid Benchmarking is based on the insight that the relative value of benchmark items depends on an LM's capability level, suggesting that evaluation should adapt to each LM. Methodologically, Fluid Benchmarking estimates an item response model based on existing LM evaluation results and uses the inferred quantities to select evaluation items dynamically, similar to computerized adaptive testing in education. In our experiments, we compare Fluid Benchmarking against the common practice of random item sampling as well as more sophisticated baselines, including alternative methods grounded in item response theory. We examine four dimensions -- efficiency, validity, variance, and saturation -- and find that Fluid Benchmarking achieves superior performance in all of them (e.g., higher validity and less variance on MMLU with fifty times fewer items). Our analysis shows that the two components of Fluid Benchmarking have distinct effects: item response theory, used to map performance into a latent ability space, increases validity, while dynamic item selection reduces variance. Overall, our results suggest that LM benchmarking can be substantially improved by moving beyond static evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。