arXiv:2604.21769cs.AIcs.CY2026-04中稿 · the 2026 ACM Confe…被引 1

让用户自定义评估标准,让大模型排行榜更贴近真实需求。

Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards

论文配图:Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
图 1 · 摘自论文原文
  • 设计交互式可视化界面,让用户按需选择和加权提示片段进行评估。
  • 发现现有榜单受数据偏斜影响,模型排名在不同提示类型中差异显著。
  • 适合关心模型实际应用场景的开发者与决策者使用。

大模型排行榜广泛用于模型比较和部署决策,但其排名由基准设计者设定的评估优先级决定,而非真实用户和组织的多样化目标与约束。单一综合得分常掩盖模型在不同提示类型和组合下的表现差异。本文深入分析了LMArena(前称Chatbot Arena)基准所用数据集,通过设计交互式可视化界面作为设计探针,揭示数据集在特定话题上高度倾斜,模型排名随提示片段变化,且偏好判断被用于超出其本意的范围。基于此分析,我们提出一种允许用户通过选择和加权提示片段来定义自身评估优先级的可视化工具,并探索相应排名变化。定性研究显示,该交互方法提升了透明度,支持更符合场景的模型评估,为未来设计和使用大模型排行榜提供了新方向。

原文摘要 · Abstract (English)

LLM leaderboards are widely used to compare models and guide deployment decisions. However, leaderboard rankings are shaped by evaluation priorities set by benchmark designers, rather than by the diverse goals and constraints of actual users and organizations. A single aggregate score often obscures how models behave across different prompt types and compositions. In this work, we conduct an in-depth analysis of the dataset used in the LMArena (formerly Chatbot Arena) benchmark and investigate this evaluation challenge by designing an interactive visualization interface as a design probe. Our analysis reveals that the dataset is heavily skewed toward certain topics, that model rankings vary across prompt slices, and that preference-based judgments are used in ways that blur their intended scope. Building on this analysis, we introduce a visualization interface that allows users to define their own evaluation priorities by selecting and weighting prompt slices and to explore how rankings change accordingly. A qualitative study suggests that this interactive approach improves transparency and supports more context-specific model evaluation, pointing toward alternative ways to design and use LLM leaderboards.

大模型评估用户定制可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。