提出新统计框架,更准评估大模型聊天机器人性能
A Statistical Framework for Ranking LLM-Based Chatbots
- 用因子化平局模型更好处理人类评分中的平局情况
- 能建模模型间性能相关性,识别出性能梯队
- 解决参数不唯一问题,确保结果稳定可解释
大语言模型(LLMs)已改变自然语言处理,Chatbot Arena等平台通过数百万次人工对比评价,为开放对话任务的模型排名提供了重要数据。本文在此基础上提出一种统计框架,针对成对比较分析中的特定挑战进行改进:首先引入因子化平局模型,显著提升对人类评分中平局现象的拟合能力;其次扩展框架以建模竞争者间的协方差,揭示性能关系并实现直观的性能分层;第三,通过引入新约束解决参数非唯一性带来的优化难题,确保参数估计稳定且可解释。通过严格评估与大量实验验证,本框架在拟合成对比较数据方面显著优于现有方法。为支持复现与应用,我们发布了开源Python包leaderbot,实现模型与分析流程。
原文摘要 · Abstract (English)
Large language models (LLMs) have transformed natural language processing, with frameworks like Chatbot Arena providing pioneering platforms for evaluating these models. By facilitating millions of pairwise comparisons based on human judgments, Chatbot Arena has become a cornerstone in LLM evaluation, offering rich datasets for ranking models in open-ended conversational tasks. Building upon this foundation, we propose a statistical framework that incorporates key advancements to address specific challenges in pairwise comparison analysis. First, we introduce a factored tie model that enhances the ability to handle ties -- an integral aspect of human-judged comparisons -- significantly improving the model's fit to observed data. Second, we extend the framework to model covariance between competitors, enabling deeper insights into performance relationships and facilitating intuitive groupings into performance tiers. Third, we resolve optimization challenges arising from parameter non-uniqueness by introducing novel constraints, ensuring stable and interpretable parameter estimation. Through rigorous evaluation and extensive experimentation, our framework demonstrates substantial improvements over existing methods in modeling pairwise comparison data. To support reproducibility and practical adoption, we release leaderbot, an open-source Python package implementing our models and analyses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。