arXiv:2412.10513cs.AIcs.LG2024-12AAAI

用理论保证提取BERT模型决策树,发现职业性别偏见。

Extracting PAC Decision Trees from Black Box Binary Classifiers: The Gender Bias Case Study on BERT-based Language Models

  • 基于PAC框架改进决策树算法,确保提取结果可信
  • 从BERT模型中提取的决策树揭示了职业性别偏见
  • 适合关注AI公平性与可解释性的研究者

决策树因其内在可解释性而广受欢迎。在可解释人工智能中,决策树可用作复杂黑箱模型的代理模型或其部分行为的近似。该方法的关键挑战在于判断提取的决策树对原模型的表征准确性以及可信度。本文研究利用概率近似正确(PAC)框架为从AI模型中提取的决策树提供理论保真度保证。基于PAC理论结果,我们改进决策树算法,在特定条件下实现PAC保证。聚焦二分类任务,实验中从基于BERT的语言模型中提取具有PAC保证的决策树。结果表明,这些模型存在职业相关的性别偏见。

原文摘要 · Abstract (English)

Decision trees are a popular machine learning method, known for their inherent explainability. In Explainable AI, decision trees can be used as surrogate models for complex black box AI models or as approximations of parts of such models. A key challenge of this approach is determining how accurately the extracted decision tree represents the original model and to what extent it can be trusted as an approximation of their behavior. In this work, we investigate the use of the Probably Approximately Correct (PAC) framework to provide a theoretical guarantee of fidelity for decision trees extracted from AI models. Based on theoretical results from the PAC framework, we adapt a decision tree algorithm to ensure a PAC guarantee under certain conditions. We focus on binary classification and conduct experiments where we extract decision trees from BERT-based language models with PAC guarantees. Our results indicate occupational gender bias in these models.

可解释性语言模型偏见检测决策树

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。