arXiv:2603.26891cs.LGcs.AI2026-03被引 3

防止模型克隆刷榜,提升生成模型排名公正性

Strategic Candidacy in Generative AI Arenas

  • 引入自评排名机制修正偏好数据偏差
  • 理论上可抵御克隆攻击,避免排名被操纵
  • 适合关注AI评估公平性的研究者与平台方

AI竞技场通过用户对模型的成对偏好来排序生成模型,是衡量模型性能的常用方式。由于偏好数据存在噪声,模型生产者可能通过提交多个相似模型(克隆)来人为提升顶级模型的排名,损害排名质量与实用性。本文首先在理论和基于真实平台数据的模拟中,证明了克隆策略在特定条件下确实能提升排名。为此,提出新机制YRWR:要求生产者对其自身模型进行排序,并用此信息校正模型质量估计。理论证明该机制近似抗克隆,即生产者无法通过多提交克隆来显著提升排名。若生产者能正确排序自身模型,整体排名准确率将提高。模拟显示,即使生产者存在误排序,该机制仍能有效抑制克隆行为并提升准确性。

原文摘要 · Abstract (English)

AI arenas, which rank generative models from pairwise preferences of users, are a popular method for measuring the relative performance of models in the course of their organic use. Because rankings are computed from noisy preferences, there is a concern that model producers can exploit this randomness by submitting many models (e.g., multiple variants of essentially the same model) and thereby artificially improve the rank of their top models. This can lead to degradations in the quality, and therefore the usefulness, of the ranking. In this paper, we begin by establishing, both theoretically and in simulations calibrated to data from the platform Arena (formerly LMArena, Chatbot Arena), conditions under which producers can benefit from submitting clones when their goal is to be ranked highly. We then propose a new mechanism for ranking models from pairwise comparisons, called You-Rank-We-Rank (YRWR). It requires that producers submit rankings over their own models and uses these rankings to correct statistical estimates of model quality. We prove that this mechanism is approximately clone-robust, in the sense that a producer cannot improve their rank much by doing anything other than submitting each of their unique models exactly once. Moreover, to the extent that model producers are able to correctly rank their own models, YRWR improves overall ranking accuracy. In further simulations, we show that indeed the mechanism is approximately clone-robust and quantify improvements to ranking accuracy, even under producer misranking.

AI评估排名机制克隆防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。