用可解释技能分析,帮用户省钱选对大模型
Trust by Design: Skill Profiles for Transparent, Cost-Aware LLM Routing
- 通过批评家式分析拆解模型输出,提取具体能力
- 构建能力矩阵并多目标优化,平衡性能与预算
- 提供自然语言解释,适合需透明决策的场景
LLM实践者如何在不浪费成本的前提下为任务选择合适模型?我们提出BELLA(基于自动技能画像的预算高效模型选择),通过可解释的基于技能的模型选择来推荐最优方案。现有基准报告整体指标,掩盖了任务所需的特定能力以及是否可用更便宜的模型替代。BELLA通过三阶段解决该问题:(1) 使用基于批评家的分析分解模型输出,提取任务所需的细粒度技能;(2) 将技能聚类为结构化能力矩阵;(3) 多目标优化,在满足预算约束的前提下最大化性能。BELLA提供自然语言理由,实现当前黑箱路由系统所缺乏的透明性。我们描述框架架构,将其置于大模型路由与评估的背景下,并以金融推理为例,展示其多样技能需求与模型间成本差异的适用性。本框架使从业者能做出有依据的成本-性能权衡。
原文摘要 · Abstract (English)
How should Large Language Model (LLM) practitioners select the right model for a task without wasting money? We introduce BELLA (Budget-Efficient LLM Selection via Automated skill-profiling), a framework that recommends optimal LLM selection for tasks through interpretable skill-based model selection. Standard benchmarks report aggregate metrics that obscure which specific capabilities a task requires and whether a cheaper model could suffice. BELLA addresses this gap through three stages: (1) decomposing LLM outputs and extract the granular skills required by using critic-based profiling, (2) clustering skills into structured capability matrices, and (3) multi-objective optimization to select the right models to maximize performance while respecting budget constraints. BELLA provides natural-language rationale for recommendations, providing transparency that current black-box routing systems lack. We describe the framework architecture, situate it within the landscape of LLM routing and evaluation, and discuss its application to financial reasoning as a representative domain exhibiting diverse skill requirements and cost-variation across models. Our framework enables practitioners to make principled and cost-performance trade-offs for deploying LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。