用瑞士轮机制动态评估大模型多维度表现,更真实反映其竞争能力。
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
- 基于胜败记录动态配对模型,模拟多轮竞赛
- 10万次蒙特卡洛仿真得出稳定预期得分
- 可区分稳健通用型与激进专精型模型
大语言模型(LLMs)的快速涌现和多样化评测基准要求从碎片化的任务指标转向综合性的竞争排名体系。现有静态评分方法难以确定跨基准的合理权重,且无法捕捉模型在连续高风险任务中的动态竞争力与脆弱性。为此,我们提出新型竞争式瑞士轮动态框架(CSD)。CSD通过基于累积胜负记录的动态配对,在一系列精选基准上模拟多轮竞赛;采用蒙特卡洛模拟($N=100,000$次迭代)计算统计稳健的预期胜场分($E[S_m]$),消除随机配对与早期运气的影响。此外,通过参数化每轮淘汰量($T_k$)进行失效敏感性分析,可刻画模型的风险偏好,从而区分稳健型通才与激进型专才。实验证明,CSD比传统聚合评分与静态成对模型提供更细致、情境感知的排名,是迈向风险感知型下一代大模型评估的关键一步。
原文摘要 · Abstract (English)
The rapid proliferation of Large Language Models (LLMs) and diverse specialized benchmarks necessitates a shift from fragmented, task-specific metrics to a holistic, competitive ranking system that effectively aggregates performance across multiple ability dimensions. Primarily using static scoring, current evaluation methods are fundamentally limited. They struggle to determine the proper mix ratio across diverse benchmarks, and critically, they fail to capture a model's dynamic competitive fitness or its vulnerability when confronted with sequential, high-stakes tasks. To address this, we introduce the novel Competitive Swiss-System Dynamics (CSD) framework. CSD simulates a multi-round, sequential contest where models are dynamically paired across a curated sequence of benchmarks based on their accumulated win-loss record. And Monte Carlo Simulation ($N=100,000$ iterations) is used to approximate the statistically robust Expected Win Score ($E[S_m]$), which eliminates the noise of random pairing and early-round luck. Furthermore, we implement a Failure Sensitivity Analysis by parameterizing the per-round elimination quantity ($T_k$), which allows us to profile models based on their risk appetite--distinguishing between robust generalists and aggressive specialists. We demonstrate that CSD provides a more nuanced and context-aware ranking than traditional aggregate scoring and static pairwise models, representing a vital step towards risk-informed, next-generation LLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。