arXiv:2512.23487cs.LGcs.AI2025-12被引 1

用系统视角解决模型选型中的能力、成本与合规平衡难题

ML Compass: Navigating Capability, Cost, and Compliance Trade-offs in AI Model Deployment

  • 将模型选择建模为能力-成本边界上的约束优化问题
  • 发现最优配置存在三类维度:合规底线、性能上限与中间值
  • 适合关注落地实效的团队,尤其在医疗等高合规场景

本文研究组织在用户效用、部署成本和合规要求共同影响下如何选择竞争性AI模型。现有能力排行榜无法直接转化为部署决策,形成能力与部署之间的鸿沟;为此,我们从系统层面出发,将模型选择与应用结果、运营约束及能力-成本前沿联系起来。提出ML Compass框架,将模型选择视为该前沿上的约束优化问题。理论上,刻画了参数化前沿下的最优配置,发现内部度量存在三类结构:部分维度固定于合规最低要求,部分饱和于最大值,其余取由前沿曲率决定的内点值。推导出预算变化、监管收紧和技术进步对各能力维度与成本的影响规律。实现上,提出四步流程:(i) 从异构模型描述中提取低维内部度量,(ii) 从能力与成本数据估计经验前沿,(iii) 从交互结果数据学习用户或任务特异性效用函数,(iv) 利用这些组件定位能力-成本配置并推荐模型。通过两个案例验证:通用对话场景使用PRISM Alignment数据集,医疗场景使用自建HealthBench数据集。在两种环境下,框架生成的推荐——以及基于预测部署价值的部署感知排行榜——与仅看能力的排名差异显著,清晰揭示能力、成本与安全之间的权衡如何影响最优模型选择。

原文摘要 · Abstract (English)

We study how organizations should select among competing AI models when user utility, deployment costs, and compliance requirements jointly matter. Widely used capability leaderboards do not translate directly into deployment decisions, creating a capability -- deployment gap; to bridge it, we take a systems-level view in which model choice is tied to application outcomes, operating constraints, and a capability-cost frontier. We develop ML Compass, a framework that treats model selection as constrained optimization over this frontier. On the theory side, we characterize optimal model configurations under a parametric frontier and show a three-regime structure in optimal internal measures: some dimensions are pinned at compliance minima, some saturate at maximum levels, and the remainder take interior values governed by frontier curvature. We derive comparative statics that quantify how budget changes, regulatory tightening, and technological progress propagate across capability dimensions and costs. On the implementation side, we propose a pipeline that (i) extracts low-dimensional internal measures from heterogeneous model descriptors, (ii) estimates an empirical frontier from capability and cost data, (iii) learns a user- or task-specific utility function from interaction outcome data, and (iv) uses these components to target capability-cost profiles and recommend models. We validate ML Compass with two case studies: a general-purpose conversational setting using the PRISM Alignment dataset and a healthcare setting using a custom dataset we build using HealthBench. In both environments, our framework produces recommendations -- and deployment-aware leaderboards based on predicted deployment value under constraints -- that can differ materially from capability-only rankings, and clarifies how trade-offs between capability, cost, and safety shape optimal model choice.

模型选型部署优化合规评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。