arXiv:2412.20086cs.LGcs.AI2024-12中稿 · ICSE24被引 15

用零阶梯度搜索实现高效黑盒公平性测试,无需模型结构信息。

MAFT: Efficient Model-Agnostic Fairness Testing for Deep Neural Networks via Zero-Order Gradient Search

  • 通过梯度估计与属性扰动实现无模型依赖的公平性检测
  • 发现公平性违规效率比现有方法高约32.58倍,效果高约14.69倍
  • 适用于大规模网络,适合实际部署中模型不可见场景

深度神经网络在各类应用中表现强大,但其公平性问题始终存在。已有部分高效白盒个体公平性测试方法,但黑盒方法发展滞后,性能远低于白盒方法。本文提出一种新型黑盒个体公平性测试方法——模型无关公平性测试(MAFT)。该方法通过轻量级的梯度估计与属性扰动,避免复杂符号执行等操作,显著提升可扩展性与适用性。实验表明,MAFT在有效性上达到顶尖白盒方法水平,同时在大模型上更具适用性。相比现有黑盒方法,其在发现公平性违规方面的效率提升约32.58倍,效果提升约14.69倍。

原文摘要 · Abstract (English)

Deep neural networks (DNNs) have shown powerful performance in various applications and are increasingly being used in decision-making systems. However, concerns about fairness in DNNs always persist. Some efficient white-box fairness testing methods about individual fairness have been proposed. Nevertheless, the development of black-box methods has stagnated, and the performance of existing methods is far behind that of white-box methods. In this paper, we propose a novel black-box individual fairness testing method called Model-Agnostic Fairness Testing (MAFT). By leveraging MAFT, practitioners can effectively identify and address discrimination in DL models, regardless of the specific algorithm or architecture employed. Our approach adopts lightweight procedures such as gradient estimation and attribute perturbation rather than non-trivial procedures like symbol execution, rendering it significantly more scalable and applicable than existing methods. We demonstrate that MAFT achieves the same effectiveness as state-of-the-art white-box methods whilst improving the applicability to large-scale networks. Compared to existing black-box approaches, our approach demonstrates distinguished performance in discovering fairness violations w.r.t effectiveness (approximately 14.69 times) and efficiency (approximately 32.58 times).

公平性测试黑盒检测零阶优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。