arXiv:2410.22368cs.SEcs.AI2024-10被引 3

提出新型模型性能评估体系,用两个数字综合衡量大模型表现。

Project MPG: towards a generalized performance benchmark for LLM capabilities

  • 用准确率和推理速度构建通用评估指标,避免依赖复杂评分系统。
  • 在多个任务上验证,与主流评测相关性更高,尤其优于MMLU榜单。
  • 适合非专家快速比较模型,也适用于细分领域性能分析。

当前大模型评估任务种类繁多,但决策者往往需要一个简洁的数值作为参考。现有方法多基于Elo评分,成本高且耗时。本文提出Project MPG(Model Performance and Goodness)——一种不依赖Elo的聚合评估方案,生成两个核心指标:‘好度’(答案准确率)和‘快度’(成本或每秒查询数)。通过横向对比模型,给出整体及子领域排名。结果显示,该方法得分与Chatbot Arena的原始皮尔逊相关性显著,甚至优于MMLU排行榜与Chatbot Arena的相关性。

原文摘要 · Abstract (English)

There exists an extremely wide array of LLM benchmarking tasks, whereas oftentimes a single number is the most actionable for decision-making, especially by non-experts. No such aggregation schema exists that is not Elo-based, which could be costly or time-consuming. Here we propose a method to aggregate performance across a general space of benchmarks, nicknamed Project "MPG," dubbed Model Performance and Goodness, additionally referencing a metric widely understood to be an important yet inaccurate and crude measure of car performance. Here, we create two numbers: a "Goodness" number (answer accuracy) and a "Fastness" number (cost or QPS). We compare models against each other and present a ranking according to our general metric as well as subdomains. We find significant agreement between the raw Pearson correlation of our scores and those of Chatbot Arena, even improving on the correlation of the MMLU leaderboard to Chatbot Arena.

大模型评估性能基准量化指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。