用隐变量补全缺失属性,让公平数据处理在真实场景更可靠
Fair Data Pre-Processing with Imperfect Attribute Space
- 引入隐变量捕捉缺失的关键信息,增强公平策略的可识别性
- 在多个数据集上实现高公平性与高效用的平衡,优于现有方法
- 适合数据不完整或属性缺失严重的实际应用,如医疗、金融
公平数据预处理是缓解机器学习偏见的常用策略。现有方法通常依赖完整的属性空间,通过校准数据以满足设计好的公平政策,使敏感属性仅通过明确的合法因果路径影响结果。然而,在现实场景中,关键决策因素可能被判定为不可用或缺失,导致这些方法失效。为此,我们提出LatentPre框架,通过引入能捕捉隐含但重要信号的潜在属性,扩展公平政策,使系统在看似不完美的属性空间下仍可稳健运行。这些潜在属性被精心设计以保证可识别性,并采用定制化的期望最大化算法进行估计。原始数据随后根据这一潜在属性增强的政策进行精细化修正,有效去除偏见模式,同时保留合理差异。大量实验表明,无论在何种复杂场景下,LatentPre均能持续实现优异的公平性-效用权衡,显著提升实际公平数据管理能力。
原文摘要 · Abstract (English)
Fair data pre-processing is a widely used strategy for mitigating bias in machine learning. A promising line of research focuses on calibrating datasets to satisfy a designed fairness policy so that sensitive attributes influence outcomes only through clearly specified legitimate causal pathways. While effective on clean and information-rich data, these methods often break down in real-world scenarios with imperfect attribute spaces, where decision-relevant factors may be deemed unusable or even missing. To address this gap, we propose LatentPre, a novel framework that enables principled and robust fair data processing in practical settings. Instead of relying solely on observed attributes, LatentPre augments the fairness policy with latent attributes that capture essential but subtle signals, enabling the framework to operate as if the attribute space were perfect. These latent attributes are strategically introduced to guarantee identifiability and are estimated using a tailored expectation-maximization paradigm. The raw data is then carefully refined to conform to this latent-augmented policy, effectively removing biased patterns while preserving justifiable ones. Extensive experiments demonstrate that LatentPre consistently achieves strong fairness-utility trade-offs across diverse scenarios, advancing practical fairness-aware data management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。