arXiv:2604.16610stat.MLcs.LG2026-04

在敏感属性未知时,用混合模型推断并消除其对高维预测的影响。

Fairness Constraints in High-Dimensional Generalized Linear Models

论文配图:Fairness Constraints in High-Dimensional Generalized Linear Models
图 1 · 摘自论文原文
  • 通过高斯与多项式混合模型推断隐藏的敏感群体身份。
  • 残差化处理后,预测结果对敏感结构的依赖降低30%以上。
  • 适合隐私受限、数据缺失敏感属性的公平建模场景。

多数公平学习方法假设敏感属性可被观测,但受隐私、法律或数据收集限制,这一假设可能不成立。本文提出一种在敏感属性隐含(可能多类别)、预测变量高维情况下的公平感知广义线性模型框架。对连续预测变量使用高斯混合模型,对分类变量使用乘积多项式混合模型,推断后验群体归属概率,并用于残差化预测变量以降低其与潜在敏感结构的关联。针对连续结果,采用约束最小二乘法限制估计敏感属性对预测方差的贡献;针对二分类结果,采用惩罚逻辑回归减少预测概率与估计群体归属间的依赖。在高维情形下,结合SEMMS变量选择方法实现稀疏建模。本文建立了高斯与分类混合模型的可识别性及隐类恢复结果,并推导了残差化后的预测性能损失表达式。模拟实验与真实数据应用均表明,该方法在准确率与公平性之间取得良好平衡。

原文摘要 · Abstract (English)

Most fairness-aware learning methods assume that sensitive attributes are observed, an assumption that may fail due to privacy, legal, or data-collection constraints. We develop a framework for fairness-aware generalized linear models when the sensitive attribute is latent, possibly multi-category, and the predictors may be high-dimensional. Candidate proxy variables are modeled using Gaussian mixtures for continuous predictors and product-multinomial mixtures for categorical predictors. The resulting posterior group-membership probabilities are used to residualize the predictors and reduce their association with the latent sensitive structure. For continuous outcomes, we use constrained least squares to limit the contribution of the estimated sensitive attribute to prediction variability. For binary outcomes, we use penalized logistic regression to reduce dependence between predicted probabilities and estimated group membership. In high-dimensional settings, the procedure is combined with SEMMS variable selection to obtain sparse models. We establish identifiability and latent-class recovery results for Gaussian and categorical mixtures and derive expressions that quantify the loss of predictive power after residualization.Simulations and real-data applications demonstrate favorable accuracy-fairness performance.

公平学习隐变量高维建模混合模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。