提出一种闭环公平决策的乐观可行搜索方法,兼顾公平与服务率约束。
Optimistic Feasible Search for Closed-Loop Fair Threshold Decision-Making
- 基于置信区间网格搜索,动态选择满足公平约束的最优阈值。
- 在合成与半真实数据上,奖励更高且约束违规累积更少。
- 适合低维可解释策略,尤其适用于信贷、司法等高风险场景。
闭环决策系统(如信贷、筛选或再犯风险评估)常面临公平性与服务率约束,且决策会改变未来样本分布,导致数据非平稳并加剧不平等。本文研究在群体平等(DP)及可选服务率约束下,从贝叶斯反馈中在线学习一维阈值策略的问题。学习者每轮仅观察一个标量评分,选择阈值;仅对所选阈值返回奖励与约束残差。我们提出乐观可行搜索(OFS),一种基于网格的简单方法,为每个候选阈值维护奖励与约束残差的置信区间。每轮选择在置信区间内可行且乐观奖励最大的阈值;若无可行阈值,则选择乐观约束违反最小的。该设计直接聚焦于可行且高收益的阈值,在低维、可解释策略类中表现尤佳。我们在(i)具有稳定收缩动力学的合成基准,以及(ii)基于德国信用与COMPAS数据构建的两个半合成基准上评估OFS。所有环境中,其奖励更高且累积约束违规更小,接近最优固定阈值的理论上限。实验可复现,输出设计支持双盲评审。
原文摘要 · Abstract (English)
Closed-loop decision-making systems (e.g., lending, screening, or recidivism risk assessment) often operate under fairness and service constraints while inducing feedback effects: decisions change who appears in the future, yielding non-stationary data and potentially amplifying disparities. We study online learning of a one-dimensional threshold policy from bandit feedback under demographic parity (DP) and, optionally, service-rate constraints. The learner observes only a scalar score each round and selects a threshold; reward and constraint residuals are revealed only for the chosen threshold. We propose Optimistic Feasible Search (OFS), a simple grid-based method that maintains confidence bounds for reward and constraint residuals for each candidate threshold. At each round, OFS selects a threshold that appears feasible under confidence bounds and, among those, maximizes optimistic reward; if no threshold appears feasible, OFS selects the threshold minimizing optimistic constraint violation. This design directly targets feasible high-utility thresholds and is particularly effective for low-dimensional, interpretable policy classes where discretization is natural. We evaluate OFS on (i) a synthetic closed-loop benchmark with stable contraction dynamics and (ii) two semi-synthetic closed-loop benchmarks grounded in German Credit and COMPAS, constructed by training a score model and feeding group-dependent acceptance decisions back into population composition. Across all environments, OFS achieves higher reward with smaller cumulative constraint violation than unconstrained and primal-dual bandit baselines, and is near-oracle relative to the best feasible fixed threshold under the same sweep procedure. Experiments are reproducible and organized with double-blind-friendly relative outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。