用三人博弈模型改进大模型评估,减少数据冗余带来的偏差
Re-evaluating Open-ended Evaluation of Large Language Models
- 将评估设计为三人博弈,避免传统评分系统对重复数据的敏感
- 新方法生成更合理的大模型排名,减少偏差影响
- 适合关注大模型公平评估的研究者与评测工程师
评估传统上聚焦于特定技能的候选模型排序。现代通用模型如大语言模型(LLMs)已显著超越这一范式。开放性评估系统通过用户提交的提示词对比候选模型,成为流行方案。尽管优势明显,我们发现当前基于Elo的评分体系易受数据冗余影响,可能放大无意或有意偏差。为此,我们提出评估作为三玩家博弈,并引入新的博弈论解概念,以增强对冗余的鲁棒性。结果表明,该方法产生直观的评分结果,揭示了大模型竞争格局的深层特征。
原文摘要 · Abstract (English)
Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Open-ended evaluation systems, where candidate models are compared on user-submitted prompts, have emerged as a popular solution. Despite their many advantages, we show that the current Elo-based rating systems can be susceptible to and even reinforce biases in data, intentional or accidental, due to their sensitivity to redundancies. To address this issue, we propose evaluation as a 3-player game, and introduce novel game-theoretic solution concepts to ensure robustness to redundancy. We show that our method leads to intuitive ratings and provide insights into the competitive landscape of LLM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。