arXiv:2507.18561cs.LGcs.AI2025-07

用合成数据解决公平性测试缺真实人群信息的难题

Beyond Internal Data: Constructing Complete Datasets for Fairness Testing

  • 通过重叠数据集构建含人口属性的合成数据
  • 合成数据上的公平性指标与真实数据一致
  • 适合需独立审计但数据受限的行业场景

随着AI在高风险领域广泛应用,公平性测试日益重要。然而,现实中因法律和隐私限制,难以获取包含人口属性的真实数据,且内部历史数据往往缺乏代表性。本文提出利用多个重叠数据集构建包含人口属性的合成数据,以准确反映受保护属性与模型特征间的关联。通过对比真实数据验证合成数据的保真度,实证表明基于合成数据计算的公平性指标与真实数据结果一致。该方法为无法获取完整数据的公平性测试提供了可行路径,支持独立、模型无关的评估,在真实数据受限时可作为有效替代。

原文摘要 · Abstract (English)

As AI becomes prevalent in high-risk domains and decision-making, it is essential to test for potential harms and biases. This urgency is reflected by the global emergence of AI regulations that emphasise fairness and adequate testing, with some mandating independent bias audits. However, procuring the necessary data for fairness testing remains a significant challenge. Particularly in industry settings, legal and privacy concerns restrict the collection of demographic data required to assess group disparities, and auditors face practical and cultural challenges in gaining access to data. Further, internal historical datasets are often insufficiently representative to identify real-world biases. This work focuses on evaluating classifier fairness when complete datasets including demographics are inaccessible. We propose leveraging separate overlapping datasets to construct complete synthetic data that includes demographic information and accurately reflects the underlying relationships between protected attributes and model features. We validate the fidelity of the synthetic data by comparing it to real data, and empirically demonstrate that fairness metrics derived from testing on such synthetic data are consistent with those obtained from real data. This work, therefore, offers a path to overcome real-world data scarcity for fairness testing, enabling independent, model-agnostic evaluation of fairness, and serving as a viable substitute where real data is limited.

公平性测试合成数据隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。