arXiv:2605.15761cs.LG2026-05

提出统一扰动框架,揭示排行榜易被小改动操控

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation

论文配图:A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation
图 1 · 摘自论文原文
  • 基于影响估计构建三类扰动分析框架
  • 不足1%的扰动即可改变冠军模型和排名一致性
  • 可高效操纵特定模型排名,适合评估者使用

评估排行榜如LMArena在衡量大语言模型性能中起核心作用,通过聚合成对人类偏好得出模型排名,但其鲁棒性仍不明确。本文提出一个统一的扰动框架,利用基于影响的近似方法,在结构化数据修改下分析Bradley-Terry排行榜。研究了三种比赛级扰动——删除、添加、翻转,以及选手移除,评估其对前k名成员、全局排名一致性(以Kendall's tau衡量)及置信区间不确定性的影响。在Chatbot Arena及六个额外的成对比较数据集上,我们发现现代排行榜在所有三个目标下均非鲁棒:小于1%的定向扰动即可改变排名第一的模型,降低Kendall's tau,并改变置信区间。除鲁棒性审计外,相同的影响得分还可实现高效定向扰动,促进或削弱特定模型表现,且所需操作少于以往操纵和主动采样基线。通过归一化的数据集级鲁棒性评分总结这些影响,该框架为审计排行榜稳定性并推动更鲁棒的评估协议提供了实用工具。

原文摘要 · Abstract (English)

Evaluation leaderboards such as LMArena play a central role in benchmarking large language models by aggregating pairwise human preferences into model rankings, yet the robustness of these rankings remains poorly understood. We present a unified perturbation framework for analyzing Bradley-Terry leaderboards under structured data modifications using influence-based approximations. Our framework studies three match-level perturbations -- Drop, Add, and Flip -- together with player removal, and evaluates their effects on top-k membership, global ranking consistency via Kendall's tau, and confidence-interval-based uncertainty. Across Chatbot Arena and six additional pairwise-comparison datasets, we show that modern leaderboards are non-robust across all three objectives: sub-1% targeted perturbations can change the top-ranked model, degrade Kendall's tau, and alter confidence intervals. Beyond robustness auditing, we show that the same influence scores enable efficient targeted perturbations, promoting or demoting specific models and reducing target-model uncertainty with fewer actions than previous manipulation and active-sampling baselines. By summarizing these effects with normalized dataset-level robustness scores, our framework provides a practical and helpful tool for auditing leaderboard stability and motivating more robust evaluation protocols.

排行榜鲁棒性模型评估扰动分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。