提出可微且满足硬约束的决策策略,提升优化模型的实用性与训练稳定性。
Smooth Learning with Hard Constraints via Legendre-Regularized Policies
- 将决策建模为原可行域上的正则化优化问题,保证可行性与可微性。
- 理论证明其映射唯一、连续且可任意光滑,雅可比矩阵显式可算。
- 适用于需要精确约束的场景,如资源分配、供应链决策等工业优化问题。
本文从策略类设计的角度重新审视上下文优化问题。理想的策略类应足够表达复杂上下文-决策关系,强制执行硬可行性约束而非软惩罚项,并保持足够平滑以支持基于梯度的下游训练。现有方法通常仅关注部分要求。本文提出勒让德正则化策略,将决策参数化为在原始可行域上求解正则化优化问题的解。该构造确保策略天然可行,且对学习的隐变量参数可微。我们证明:对应优化器映射是单值的,映射至可行集的相对内部,具有显式雅可比,满足利普希茨连续性,且可任意光滑。同时建立通用逼近结果,表明该类策略可在紧致上下文集上逼近任意连续可行策略。该框架统一了显式正则化优化器与隐式扰动平滑优化器。在上下文新闻商和资源分配问题上的实验表明,本方法在预测性能上优于基准方法。
原文摘要 · Abstract (English)
We revisit contextual optimization from the perspective of policy class design. A desirable policy class should be expressive enough to learn rich context-decision relationships, should enforce hard feasibility constraints rather than soft penalty terms, and should remain smooth enough for gradient-based training on downstream decision losses. Existing approaches usually emphasize only part of these requirements. We propose Legendre-regularized policies, which parameterize decisions as solutions of regularized optimization problems over the original feasible region. This construction yields policies that are feasible by construction and differentiable with respect to learned latent parameters. We prove that the associated optimizer map is single-valued, maps onto the relative interior of the feasible set, admits an explicit Jacobian, is Lipschitz continuous, and can be made arbitrarily smooth. We also establish a universal approximation result showing that the proposed class can approximate any continuous feasible policy on compact context sets. The framework unifies explicitly regularized optimizers and implicit perturbation-based smooth optimizers. Experiments on contextual newsvendor and resource allocation problems show that our approach improves prescriptive performance relative to the benchmark methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。