arXiv:2603.15220cs.AI2026-03

用插值偏好学习破解大模型排行榜匿名性,精准识别模型身份

InterPol: De-anonymizing LM Arena via Interpolated Preference Learning

  • 通过模型插值生成难样本,捕捉深层风格特征
  • 在真实对抗数据上识别准确率显著超越现有方法
  • 揭示排行榜匿名机制漏洞,适合安全与评测研究者参考

严格匿名的模型响应是基于投票的排行榜(如LM Arena)可靠性的关键。尽管已有研究尝试使用TF-IDF或词袋等简单统计特征来破坏这一假设,但这些方法往往难以区分风格相似或同家族的模型。为克服上述局限并揭示潜在风险的严重性,本文提出INTERPOL——一种基于模型驱动的识别框架,利用插值偏好数据学习区分目标模型与其他模型的能力。具体而言,INTERPOL通过模型插值合成硬负样本,并采用自适应课程学习策略,捕捉表面统计特征所忽略的深层风格模式。大量实验表明,INTERPOL在识别准确率上显著优于现有基线。此外,我们通过在Arena对战数据上的排名操纵模拟,量化了该发现的实际威胁。

原文摘要 · Abstract (English)

Strict anonymity of model responses is a key for the reliability of voting-based leaderboards, such as LM Arena. While prior studies have attempted to compromise this assumption using simple statistical features like TF-IDF or bag-ofwords, these methods often lack the discriminative power to distinguish between stylistically similar or within-family models. To overcome these limitations and expose the severity of vulnerability, we introduce INTERPOL, a model-driven identification framework that learns to distinguish target models from others using interpolated preference data. Specifically, INTERPOL captures deep stylistic patterns that superficial statistical features miss by synthesizing hard negative samples through model interpolation and employing an adaptive curriculum learning strategy. Extensive experiments demonstrate that INTERPOL significantly outperforms existing baselines in identification accuracy. Furthermore, we quantify the real-world threat of our findings through ranking manipulation simulations on Arena battle data.

模型识别隐私安全排行榜评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。