端到端学习资源分配策略,突破传统方法的容量约束瓶颈。
Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints

- 直接优化决策价值,通过双重价格反向传播训练模型
- 在六组数据中均优于基线,容量紧约束下延迟显著降低
- 适合资源稀缺且需保证可行性的真实场景
许多社会服务需按序为陆续到来的个体分配有限资源(如住房援助或医院干预),每一步必须即时决策,且长期资源使用量不得超过容量。本文研究如何从观测数据中学习此类分配策略。传统方法为“决策盲”:先对每个选项单独建模预测结果,再根据模型估计资源价格,最终选择预测收益减去价格最大的选项。本文提出端到端训练:将部署策略的离线价值估计值对双重价格进行反向传播,直接优化整体策略。研究两种形式:一种精确但非凸,另一种为凸松弛,其最优解期望满足容量约束,且次优性不超过平滑温度的线性项与臂数对数项之和。所有方法在资源按容量速率补充的排队模拟中测试。在六个数据集上,两种端到端方法在不同延迟成本下均排名第一,包括零延迟成本;当容量受限时,传统基线常违反约束,导致更长排队延迟。在最大规模数据集(7万例患者)中,端到端方法仍显著提升策略价值,该优势在容量匹配的神经网络基线中依然存在。若真实标签可测,传统回归仍是更强的预测器;但面对真正稀缺资源且可行性关键时,端到端方法更优。
原文摘要 · Abstract (English)
Many social services assign scarce resources, such as housing assistance or hospital interventions, to people who arrive one at a time: each arrival must receive a decision immediately, and the long-run usage of every resource must stay within its capacity. We study how to learn such an assignment policy from logged observational data. The standard pipeline is decision-blind: fit one outcome model per arm by regression, price each capacitated resource from the fitted models, and assign each arrival the arm whose predicted outcome minus price is largest. We instead train the outcome models end-to-end, differentiating an off-policy estimate of the deployed policy's value through the dual prices themselves. We study two formulations: an exact nonconvex one, and a convex relaxation whose optimum always satisfies the capacity constraints in expectation and which is suboptimal by at most a term linear in the smoothing temperature and logarithmic in the number of arms. Every method is evaluated in a queueing simulation with resources replenished at their capacity rates. Across six datasets, the two end-to-end variants take the top slots on a deployment-adjusted value index at every delay cost, including zero; when capacities are binding, decision-blind baselines frequently violate them and incur much longer queueing delays. On the largest dataset, a hospital cohort of seventy thousand patients, end-to-end training also achieves significantly higher policy value, a margin that survives a capacity-matched neural baseline. Flexible decision-blind regression remains the stronger pure predictor where ground truth is measurable; end-to-end training is best suited to settings where resources are genuinely scarce and feasibility matters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。