首个开放平台,用于全面对比大模型路由选择工具。
RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
- 构建覆盖广泛知识领域的评测数据集,分难度层级
- 支持多维度指标自动评估,生成可更新排行榜
- 适合研究者和开发者选型,推动路由工具标准化
当前大模型生态涵盖多种规模、能力与成本的模型,单一模型无法适应所有场景,因此模型路由成为关键。但路由工具快速涌现,选择困难。为此,我们提出 RouterArena——首个开放平台,实现大模型路由的全面比较。该平台具备:(1)覆盖广泛知识领域的结构化数据集;(2)各领域区分难度等级;(3)丰富的评估指标列表;(4)自动化排行榜更新框架。基于此框架,我们已生成初始排行榜,包含详细指标对比(见图1)。新路由评估框架开源于 https://github.com/RouteWorks/RouterArena,排行榜地址为 https://routeworks.github.io/。
原文摘要 · Abstract (English)
Today's LLM ecosystem comprises a wide spectrum of models that differ in size, capability, and cost. No single model is optimal for all scenarios; hence, LLM routers have become essential for selecting the most appropriate model under varying circumstances. However, the rapid emergence of various routers makes choosing the right one increasingly challenging. To address this problem, we need a comprehensive router comparison and a standardized leaderboard, similar to those available for models. In this work, we introduce RouterArena, the first open platform enabling comprehensive comparison of LLM routers. RouterArena has (1) a principally constructed dataset with broad knowledge domain coverage, (2) distinguishable difficulty levels for each domain, (3) an extensive list of evaluation metrics, and (4) an automated framework for leaderboard updates. Leveraging our framework, we have produced the initial leaderboard with detailed metrics comparison as shown in Figure 1. Our framework for evaluating new routers is on https://github.com/RouteWorks/RouterArena. Our leaderboard is on https://routeworks.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。