提出新评估框架,提升大模型路由系统公平性与鲁棒性
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
- 用隐藏状态捕捉模型不确定性,替代传统概率输出
- 新方法在准确率场景下相对提升16.68%~18.86%
- 适用于不同模型规模、任务类型和智能体工作流
大型语言模型(LLMs)虽取得成功,但成本与隐私限制要求将小型模型本地部署,复杂查询仍交由云端模型处理。现有路由器评估体系不系统,忽视场景适配性和分布外鲁棒性。本文提出RouterXBench评估框架,涵盖路由器能力、场景对齐性与跨域鲁棒性三个维度。不同于依赖输出概率或外部嵌入的方法,我们利用生成答案前的内部隐藏状态,以捕捉模型不确定性。提出ProbeDirichlet:通过可学习的狄利克雷分布聚合多层隐藏状态,采用概率化训练。在多领域数据上训练后,该方法在同分布与分布外场景中均表现稳健。实验表明,其在路由器能力与高精度场景下分别较最优基线提升16.68%与18.86%,且在不同模型家族、规模、异构任务及智能体工作流中保持稳定性能。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved success, but cost and privacy constraints necessitate deploying smaller models locally while offloading complex queries to cloud-based models. Existing router evaluations are unsystematic, overlooking scenario-specific requirements and out-of-distribution robustness. We propose RouterXBench, a principled evaluation framework with three dimensions: router ability, scenario alignment, and cross-domain robustness. Unlike prior work that relies on output probabilities or external embeddings, we utilize internal hidden states that capture model uncertainty before answer generation. We introduce ProbeDirichlet, a lightweight router that aggregates cross-layer hidden states via learnable Dirichlet distributions with probabilistic training. Trained on multi-domain data, it generalizes robustly across in-domain and out-of-distribution scenarios. Our results show ProbeDirichlet achieves 16.68% and 18.86% relative improvements over the best baselines in router ability and high-accuracy scenarios, with consistent performance across model families, model scales, heterogeneous tasks, and agentic workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。