arXiv:2608.29560cs.LGmath.OC2026-09

在预算有限时,如何选对模型处理任务?关键在于评估结果的不确定性。

Which LLM for Which Work? Budgeted Model Allocation under Uncertain Evaluation

论文配图:Which LLM for Which Work? Budgeted Model Allocation under Uncertain Evaluation
图 1 · 摘自论文原文
  • 通过因果实验动态聚焦评估最需验证的模型-任务组合
  • 发现多数情况下现有证据无法确定最优分配方案
  • 即使反复评估,随机性仍导致质量表始终不确定

企业在固定AI预算下需决定每项重复性任务由哪个大语言模型承担。难题在于缺乏各模型在各项任务上的性能对比表。若拥有该表,决策即为多选背包问题,可轻松求解;但估计该表存在双重困难:模型很少在同一任务上被比较,且记录得分常为代理指标而非企业真正关心的结果。因果与离策略方法修复第一类问题但依赖第二类假设,评估者验证法估计第二类却无法直接支持决策。更严重的是,增加重评无法解决第二类问题:评分随机性决定哪些请求被评分,而非评分本身生成方式,因此无论投入多少评估,表格始终存在不确定性。然而,即便表格不确定,部署决策仍可能被确定。我们提出检验:是否存在一个分配方案,在所有与证据一致的质量表中都最优?对于固定预算问题,该检验可通过两次求解实现:一次基于估计表,一次基于最不利表。结果一致则认证该方案;不一致则标识出需进一步证据的关键模型-任务对。我们提出CASE(因果主动序列实验),针对这些关键对进行评估并随证据积累重复测试。在真实生产日志中,测量缺陷是主要损失来源:即使精确修正分配,仍保留大部分损失;随机重评无法消除该损失。实验表明,现有证据常不足以确定最优分配。在付费软件任务中,提升模型质量信息比在相同估计下优化分配带来更多节省。

原文摘要 · Abstract (English)

A company with a fixed artificial intelligence (AI) budget must decide which large language model (LLM) handles each recurring workload. What it lacks is the quality table, how well each model performs on each workload. Given that table, the decision is a multiple-choice knapsack problem and is routine to solve, so estimating it is the difficulty, and that estimation fails in two ways. Models are rarely compared on the same work, and the recorded score is usually a proxy rather than the outcome the company values. Causal and off-policy methods repair the first but condition on the second, while evaluator-validation methods estimate the second but stop short of the decision. Worse, buying more re-evaluation cannot settle the second: randomization governs which requests are scored, not how a score is produced, so the table stays uncertain however much evaluation is purchased. Yet the deployment decision may still be determined even when the table is not. We therefore ask whether one assignment stays optimal across every quality table consistent with the evidence. For the fixed-budget problem, this admits an exact two-solve certificate: solve once at the estimated table and once at a least-favourable table. Agreement certifies the assignment; disagreement identifies the model-workload pairs where further evidence can matter. We propose CASE (causal active sequential experimentation), which targets evaluation to those pairs and repeats the test as evidence accumulates. On a production log, the measurement failure is the larger of the two: correcting assignment exactly still leaves most of the loss, and randomized re-evaluation does not remove it. In our experiments, the available evidence often does not determine the assignment. On paid software tasks, better information about model quality yields more savings than further optimization of the assignment on the same estimates.

模型选择因果推断预算优化不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。