用市场机制动态定价,让公司高效分配昂贵的机器学习训练资源。
Quota Marketplace: Dynamic Pricing for Efficient Allocation of ML Training Resources

- 引入价值感知的市场机制,用户可报价表达任务重要性
- 在异构需求下仍保证帕累托效率和最大最小公平性
- 已在谷歌落地,显著提升资源利用率与组织优先级对齐
近年来机器学习训练资源需求激增,供需矛盾突出。现有资源分配机制(如Karma)虽能保障帕累托效率和最大最小公平性,但在面对不同工作负载价值差异较大的场景时表现不佳。本文提出并实现了Quota Marketplace——一种基于市场的芯片(如GPU)资源分配机制,专门应对异构价值需求。该机制在谷歌内部部署,通过允许用户表达任务价值并根据供需动态定价,实现资源分配与组织目标一致。理论分析证明其可维持关键公平性属性,实际数据显示其有效提升了资源利用效率,解锁了多项业务收益。
原文摘要 · Abstract (English)
The escalating demand for Machine Learning (ML) training resources in recent years has resulted in a substantial gap between the high demand and the available supply. Efficient allocation of these scarce and expensive resources is crucial for organizations to maximize their return on investment. Existing resource allocation mechanisms, like Karma [OSDI'23], are designed to guarantee Pareto efficiency and max-min fairness in settings with dynamic (time-varying) user demands, but fail to preserve these key properties in the presence of demands with heterogeneous values. Given the ubiquity and inevitability of heterogeneity in organizational values of different workloads, effective resource allocation policies must accommodate these variations. In this paper, we describe the design, implementation, deployment, and theoretical analysis of Quota Marketplace, a market-based mechanism to efficiently allocate ML training chips (like GPUs), explicitly addressing scenarios with demands of heterogeneous value. We detail the implementation of this mechanism within Google and present metrics that demonstrate its impact. We also discuss many business-critical requirements that the Quota Marketplace handles quite effectively, and document the gains and opportunities it has unlocked. We establish theoretically how this market-based approach achieves the essential properties of Pareto efficiency and max-min fairness by allowing the users to express the value of their workloads and enabling dynamic resource pricing based on supply and demand fluctuations. Ultimately, the market facilitates resource allocation that aligns with organizational priorities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。