用数据预处理让强化学习决策更公平,避免医疗资源分配偏袒特定群体。
Counterfactually Fair Reinforcement Learning via Sequential Data Preprocessing
- 通过序列化数据预处理,实现因果公平的多阶段决策。
- 在模拟中同时降低不公平性并保持最优治疗效果。
- 适合关注医疗公平性与强化学习结合的研究者。
在医疗领域应用强化学习时,需动态匹配干预措施以最大化群体效益。然而,学习到的策略可能过度向某一群体分配有效干预,加剧弱势群体间的不平等。此类偏差常见于多阶段决策过程,且具有自我强化特性,若不解决,可能导致护理或治疗权益受限。因果公平(CF)基于因果推断,为公平性建模提供了有力工具。本文提出一种通用的公平序贯决策框架,理论刻画了最优CF策略并证明其平稳性,极大简化了最优策略搜索,可直接利用现有强化学习算法。该理论还启发了一种序列数据预处理算法,在加性噪声假设下实现CF决策。我们在模拟中验证了该方法能有效控制不公平性并达到最优价值。对一个旨在减少阿片类药物滥用的数字健康数据集分析显示,本方法显著提升了咨询资源的公平可及性。
原文摘要 · Abstract (English)
When applied in healthcare, reinforcement learning (RL) seeks to dynamically match the right interventions to subjects to maximize population benefit. However, the learned policy may disproportionately allocate efficacious actions to one subpopulation, creating or exacerbating disparities in other socioeconomically-disadvantaged subgroups. These biases tend to occur in multi-stage decision making and can be self-perpetuating, which if unaccounted for could cause serious unintended consequences that limit access to care or treatment benefit. Counterfactual fairness (CF) offers a promising statistical tool grounded in causal inference to formulate and study fairness. In this paper, we propose a general framework for fair sequential decision making. We theoretically characterize the optimal CF policy and prove its stationarity, which greatly simplifies the search for optimal CF policies by leveraging existing RL algorithms. The theory also motivates a sequential data preprocessing algorithm to achieve CF decision making under an additive noise assumption. We prove and then validate our policy learning approach in controlling unfairness and attaining optimal value through simulations. Analysis of a digital health dataset designed to reduce opioid misuse shows that our proposal greatly enhances fair access to counseling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。