删掉极少量偏好数据,顶尖大模型排名就可能翻转。
Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings
- 用快速易行的方法检测大模型排名对少数偏好数据的敏感度。
- 仅删除0.003%的人类偏好,Chatbot Arena上的榜首模型就可能改变。
- 专家标注的MT-bench数据生成的排名更稳定,适合可信评估。
我们提出一种方法,评估主流大语言模型排名系统(基于Bradley-Terry模型变体)对剔除极小比例最差偏好数据的鲁棒性。该方法计算高效且易于应用。在Chatbot Arena及其衍生平台的对决数据上应用后发现,顶级模型的排名对少量偏好数据的移除极为敏感:例如,仅剔除0.003%的人类偏好,即可导致排行榜榜首变更。我们的鲁棒性检测可定位引发排名变化的关键偏好,便于人工审查。对比发现,基于MT-bench的排名显著更稳健,可能因其使用专家标注者和精心设计的提示。此外,基于众包人类评估或大模型作为裁判的排名,在敏感性上并无系统性差异。
原文摘要 · Abstract (English)
We propose a method for evaluating the robustness of widely used LLM ranking systems -- variants of a Bradley--Terry model -- to dropping a worst-case very small fraction of preference data. Our approach is computationally fast and easy to adopt. When we apply our method to matchups from popular LLM ranking platforms, including Chatbot Arena and derivatives, we find that the rankings of top-performing models can be remarkably sensitive to the removal of a small fraction of preferences; for instance, dropping just 0.003% of human preferences can change the top-ranked model on Chatbot Arena. Our robustness check identifies the specific preferences most responsible for such ranking flips, allowing for inspection of these influential preferences. We observe that the rankings derived from MT-bench preferences are notably more robust than those from Chatbot Arena, likely due to MT-bench's use of expert annotators and carefully constructed prompts. Finally, we find that neither rankings based on crowdsourced human evaluations nor those based on LLM-as-a-judge preferences are systematically more sensitive than the other.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。