arXiv:2501.18243stat.APcs.CL2025-01被引 2

为大模型系统性能评估提供统计算法与可视化工具

Statistical multi-metric evaluation and visualization of LLM system predictive performance

  • 自动执行多指标跨数据集的统计检验
  • 支持在跨任务、跨指标下判断性能差异显著性
  • 适合模型选型与系统优化决策参考

生成式或判别式大语言模型(LLM)系统评估通常是一个多维度复杂问题。通常需在多个基准数据集上,针对一个或多个评估指标,对比多种系统配置。我们希望用统计显著性方法判断系统在单个数据集上的单个指标表现是否不同,或在单个数据集上多个指标的综合表现是否不同,或在不同数据集间的整体表现是否存在差异。此类评估可用于支持决策,例如判断某个系统组件变更(如模型选择或超参数调整)是否显著提升性能,或固定一组系统配置(如排行榜)在关注指标上是否有显著差异。本文提出一个框架实现,可自动执行正确的统计检验,合理聚合跨指标和跨数据集的统计结果(这一过程非平凡),并实现可视化。该框架在多语言代码生成基准 CrossCodeEval 上对多个前沿 LLM 进行了演示。

原文摘要 · Abstract (English)

The evaluation of generative or discriminative large language model (LLM)-based systems is often a complex multi-dimensional problem. Typically, a set of system configuration alternatives are evaluated on one or more benchmark datasets, each with one or more evaluation metrics, which may differ between datasets. We often want to evaluate -- with a statistical measure of significance -- whether systems perform differently either on a given dataset according to a single metric, on aggregate across metrics on a dataset, or across datasets. Such evaluations can be done to support decision-making, such as deciding whether a particular system component change (e.g., choice of LLM or hyperparameter values) significantly improves performance over the current system configuration, or, more generally, whether a fixed set of system configurations (e.g., a leaderboard list) have significantly different performances according to metrics of interest. We present a framework implementation that automatically performs the correct statistical tests, properly aggregates the statistical results across metrics and datasets (a nontrivial task), and can visualize the results. The framework is demonstrated on the multi-lingual code generation benchmark CrossCodeEval, for several state-of-the-art LLMs.

大模型评估统计检验多指标分析可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。