用不确定性估计优化大模型路由,兼顾成本与人类偏好。
Leveraging Uncertainty Estimation for Efficient LLM Routing
- 基于模型不确定性动态选择最合适的LLM处理请求。
- 在MT-Bench等数据集上同时提升响应质量和成本效率。
- 首次用LLM做裁判模拟人类评分,客观评估回复质量。
在边缘-云环境中部署大语言模型(LLMs)需要高效的路由策略以平衡成本与响应质量。传统方法通常依赖人工偏好数据或基准测试的准确率作为路由依据,但存在僵化和主观性问题。现有框架多关注准确率与成本,忽视了从人类偏好角度衡量响应质量。本文提出一种新型的置信度驱动的LLM路由器(Confidence-Driven LLM Router),利用不确定性估计优化路由决策。为全面评估性能,我们同时考察系统成本效率和响应质量。特别地,引入LLM-as-a-Judge来模拟人类评分,首次系统性地对比不同路由策略下的响应质量。在MT-Bench、GSM8K和MMLU上的大量实验表明,该方法优于现有先进路由方案,在保持成本效率的同时显著提升响应质量。
原文摘要 · Abstract (English)
Deploying large language models (LLMs) in edge-cloud environments requires an efficient routing strategy to balance cost and response quality. Traditional approaches prioritize either human-preference data or accuracy metrics from benchmark datasets as routing criteria, but these methods suffer from rigidity and subjectivity. Moreover, existing routing frameworks primarily focus on accuracy and cost, neglecting response quality from a human preference perspective. In this work, we propose the Confidence-Driven LLM Router, a novel framework that leverages uncertainty estimation to optimize routing decisions. To comprehensively assess routing performance, we evaluate both system cost efficiency and response quality. In particular, we introduce the novel use of LLM-as-a-Judge to simulate human rating preferences, providing the first systematic assessment of response quality across different routing strategies. Extensive experiments on MT-Bench, GSM8K, and MMLU demonstrate that our approach outperforms state-of-the-art routing methods, achieving superior response quality while maintaining cost efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。