arXiv:2503.18826cs.LGcs.AI2025-03被引 6

提出可解释的拒判机制,让分类器同时根据不确定性和不公平性拒绝预测。

Interpretable and Fair Mechanisms for Abstaining Classifiers

  • 基于不确定性和不公平性双重标准拒判,提升公平性。
  • 显著降低不同群体在通过数据上的错误率和正类决策率差异。
  • 使用规则化公平检测与情境测试,结果可被人类审查,适合高风险场景。

拒判分类器可在难以分类的样本上选择不提供预测,以权衡整体性能与预测数量。然而,现有方法常导致多数群体错误率下降而少数群体差距扩大,引发公平性问题。本文提出可解释且公平的拒判分类器 IFAC,不仅依据预测不确定性拒判,还基于不公平性进行拒绝。该方法通过规则化公平检查与情境测试实现可解释性,有效减少非拒判数据中各人口群体间的错误率与正类决策率差异。由于拒判逻辑透明可追溯,符合近期人工智能监管要求,可支持人工专家介入审查,降低歧视风险,适用于高风险决策任务。

原文摘要 · Abstract (English)

Abstaining classifiers have the option to refrain from providing a prediction for instances that are difficult to classify. The abstention mechanism is designed to trade off the classifier's performance on the accepted data while ensuring a minimum number of predictions. In this setting, often fairness concerns arise when the abstention mechanism solely reduces errors for the majority groups of the data, resulting in increased performance differences across demographic groups. While there exist a bunch of methods that aim to reduce discrimination when abstaining, there is no mechanism that can do so in an explainable way. In this paper, we fill this gap by introducing Interpretable and Fair Abstaining Classifier IFAC, an algorithm that can reject predictions both based on their uncertainty and their unfairness. By rejecting possibly unfair predictions, our method reduces error and positive decision rate differences across demographic groups of the non-rejected data. Since the unfairness-based rejections are based on an interpretable-by-design method, i.e., rule-based fairness checks and situation testing, we create a transparent process that can empower human decision-makers to review the unfair predictions and make more just decisions for them. This explainable aspect is especially important in light of recent AI regulations, mandating that any high-risk decision task should be overseen by human experts to reduce discrimination risks.

公平性可解释性拒判分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。