arXiv:2508.13040cs.LG2025-08

利用不完整数据推断模型公平性边界,解决隐私受限下的审计难题。

Beyond Internal Data: Bounding and Estimating Fairness from Incomplete Data

  • 基于内外部数据分离的特性,构建可行联合分布集
  • 可获得公平性指标的可靠上下界与估计值
  • 适合隐私受限场景下的实际公平性评估

确保AI系统公平性至关重要,尤其在信贷、招聘和医疗等高风险领域。全球新规正要求进行公平性评估与独立偏见审计。然而,获取完整的公平性测试数据仍面临重大挑战:行业实践中,法律与隐私顾虑限制了人口属性数据的收集,审计人员也常因实际与文化障碍难以访问数据。现实中,公平性测试数据往往分散于不同来源:机构持有的内部数据包含预测特征,而外部公开数据(如人口普查)则含受保护属性,二者仅提供部分边际信息。本文旨在利用这些分散数据,在无法获取完整数据时估计模型公平性。我们提出通过已有分离数据估计一组可行的联合分布,并据此计算合理的公平性指标范围。通过仿真与真实实验,证明该方法能有效获得公平性指标的有意义边界,并得到对真实指标的可靠估计。结果表明,此方法可在数据受限的实际场景中作为可行且高效的公平性测试方案。

原文摘要 · Abstract (English)

Ensuring fairness in AI systems is critical, especially in high-stakes domains such as lending, hiring, and healthcare. This urgency is reflected in emerging global regulations that mandate fairness assessments and independent bias audits. However, procuring the necessary complete data for fairness testing remains a significant challenge. In industry settings, legal and privacy concerns restrict the collection of demographic data required to assess group disparities, and auditors face practical and cultural challenges in gaining access to data. In practice, data relevant for fairness testing is often split across separate sources: internal datasets held by institutions with predictive attributes, and external public datasets such as census data containing protected attributes, each providing only partial, marginal information. Our work seeks to leverage such available separate data to estimate model fairness when complete data is inaccessible. We propose utilising the available separate data to estimate a set of feasible joint distributions and then compute the set plausible fairness metrics. Through simulation and real experiments, we demonstrate that we can derive meaningful bounds on fairness metrics and obtain reliable estimates of the true metric. Our results demonstrate that this approach can serve as a practical and effective solution for fairness testing in real-world settings where access to complete data is restricted.

公平性数据隐私偏见审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。