提出高效量化树模型公平性偏差的方法,可精确估算偏差比例与发生区域。
Quantitative Verification of Fairness in Tree Ensembles
- 基于树结构特性设计任意精度上下界估计方法
- 在5个数据集上显著优于现有测试技术,偏差识别更全面
- 适合需要精细诊断算法偏见的AI安全与合规团队
本文聚焦于树集成模型的公平性定量验证。与传统方法仅返回单个反例不同,定量验证能估计所有反例的比例并定位其出现区域,对诊断和缓解偏差至关重要。目前该方法主要应用于深度神经网络,代表性工作如DeepGemini和FairQuant均基于反例引导抽象精化框架。我们将其扩展为模型无关形式,但发现存在两个局限:(i) 仅能提供下界,(ii) 性能随规模恶化。针对树集成的离散结构特性,本文提出一种高效量化技术,可提供任意时间的上下界。在五个常用数据集上的实验表明,该方法在公平性测试中显著优于现有最先进技术。
原文摘要 · Abstract (English)
This work focuses on quantitative verification of fairness in tree ensembles. Unlike traditional verification approaches that merely return a single counterexample when the fairness is violated, quantitative verification estimates the ratio of all counterexamples and characterizes the regions where they occur, which is important information for diagnosing and mitigating bias. To date, quantitative verification has been explored almost exclusively for deep neural networks (DNNs). Representative methods, such as DeepGemini and FairQuant, all build on the core idea of Counterexample-Guided Abstraction Refinement, a generic framework that could be adapted to other model classes. We extended the framework into a model-agnostic form, but discovered two limitations: (i) it can provide only lower bounds, and (ii) its performance scales poorly. Exploiting the discrete structure of tree ensembles, our work proposes an efficient quantification technique that delivers any-time upper and lower bounds. Experiments on five widely used datasets demonstrate its effectiveness and efficiency. When applied to fairness testing, our quantification method significantly outperforms state-of-the-art testing techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。