提出路由大模型评估基准,发现路由能力随候选模型增多而显著提升。
RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs
- 构建基于8500+模型的路由评估基准,覆盖12个主流任务
- 实验证明路由性能随候选模型数量增加而提升,可超越单个最优模型
- 开源数据与工具,助力路由算法研究
路由大型语言模型(LLMs)是一种新范式,通过路由器从候选模型池中为给定输入推荐最佳模型。本文对超过8500个不同模型的全面分析揭示了路由LLMs中一种新型的模型级扩展现象:随着候选模型数量增加,强大路由器能显著提升该范式性能,甚至超越池中最佳单个模型及许多现有强模型,证实其高度前景。然而,缺乏全面且开源的路由LLM评估基准,严重阻碍了路由器的发展。为此,本文引入RouterEval,一个专为路由器研究设计的基准,包含基于超过8500个不同模型的12个流行任务(如常识推理、语义理解等)的2亿多个性能记录。利用RouterEval,对现有路由方法的广泛评估显示,多数仍存在巨大改进空间。
原文摘要 · Abstract (English)
Routing large language models (LLMs) is a new paradigm that uses a router to recommend the best LLM from a pool of candidates for a given input. In this paper, our comprehensive analysis with more than 8,500 LLMs reveals a novel model-level scaling up phenomenon in Routing LLMs, i.e., a capable router can significantly enhance the performance of this paradigm as the number of candidates increases. This improvement can even surpass the performance of the best single model in the pool and many existing strong LLMs, confirming it a highly promising paradigm. However, the lack of comprehensive and open-source benchmarks for Routing LLMs has hindered the development of routers. In this paper, we introduce RouterEval, a benchmark tailored for router research, which includes over 200,000,000 performance records for 12 popular LLM evaluations across various areas such as commonsense reasoning, semantic understanding, etc., based on over 8,500 various LLMs. Using RouterEval, extensive evaluations of existing Routing LLM methods reveal that most still have significant room for improvement. See https://github.com/MilkThink-Lab/RouterEval for all data, code and tutorial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。