研究者虽不信榜单,却仍用它做参考,关键靠同行推荐。
The Trust Paradox: How CS Researchers Engage LLM Leaderboards
- 研究者普遍怀疑榜单可靠性,但仍将其作为决策辅助。
- 同行网络是主要模型选择方式,人评榜单比静态榜单更受青睐。
- 七成研究者呼吁公开成本信息,推动评价体系改进。
大型语言模型(LLM)排行榜通过标准化基准对模型进行排名,在计算机科学领域广受关注,尽管其可靠性和鲁棒性存在已知问题。然而,这些榜单如何影响研究者的实际工作尚无实证研究。本文通过对四个子领域中八位研究者的半结构化访谈,采用反思性主题分析法,发现了一种普遍存在的实用怀疑悖论:尽管受访者普遍不信任排行榜的排名结果,但仍将其作为粗略决策依据。同行网络成为最主要的模型选择机制,而基于人类投票的排行榜始终优于静态基准排行榜。排行榜影响力在不同子领域间差异显著,表明学科文化而非个人态度决定参与度;例如,自然语言处理(NLP)研究者面临与顶尖水平对比的压力,而人机交互(HCI)和系统/隐私领域的研究者则无此压力。尽管存在差异,所有受访者均一致要求增加成本透明度(七位提及)。我们据此提出具体设计建议,如任务细分得分、成本集成、投票者人口统计披露,以使评估基础设施更贴合研究者的真实使用习惯。
原文摘要 · Abstract (English)
Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite known limitations in their reliability and robustness. Yet how they shape researchers' actual practice remains empirically uncharted. We address this gap through semi-structured interviews with eight researchers across four computer science subfields, analyzed using reflexive thematic analysis. We find a near-universal paradox of pragmatic skepticism: while participants expressed deep distrust of leaderboard rankings, they continued to use them as rough decision-making aids. Peer networks, not leaderboards, emerged as the primary model selection mechanism, and arena-based (human-voting) leaderboards were consistently preferred over static benchmark leaderboards. Leaderboard influence varied sharply across subfields, revealing that disciplinary culture, not individual attitudes, mediates engagement; for instance, NLP researchers faced state-of-the-art comparison pressure while HCI and Systems/Privacy researchers reported none. Across these differences, however, participants converged on cost transparency as the most demanded missing feature (seven of eight). We translate these findings into concrete design recommendations that align evaluation infrastructure with how researchers actually use it, such as task-specific score breakdowns, cost integration, and voter-demographic disclosure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。