arXiv:2502.16380cs.LGcs.AI2025-02ICML被引 1

提出新方法识别模型中无法改变预测结果的群体,可提前发现潜在不公平。

Understanding Fixed Predictions via Confined Regions

  • 通过寻找特征空间中的封闭区域定位固定预测
  • 在无代表数据集时仍能验证未来个体的可回溯性
  • 适合关注算法公平性与可解释性的研究者

机器学习模型可能对个体给出无法更改的预测结果。现有审计方法依赖具体数据点,需已有数据集且难以预判外部样本中的固定预测。本文提出新范式:通过识别特征空间中所有个体均获得固定预测的封闭区域,实现对外部数据的可回溯性认证。该方法适用于缺乏代表性数据集的场景,并提供可解释的固定预测描述。我们开发了基于混合整数二次约束规划的快速算法,用于线性分类器的封闭区域发现。在多个应用场景中开展全面实证研究,结果表明,传统点级验证方法无法预判未来个体的固定预测,而本方法不仅能识别,还能提供清晰解释。

原文摘要 · Abstract (English)

Machine learning models can assign fixed predictions that preclude individuals from changing their outcome. Existing approaches to audit fixed predictions do so on a pointwise basis, which requires access to an existing dataset of individuals and may fail to anticipate fixed predictions in out-of-sample data. This work presents a new paradigm to identify fixed predictions by finding confined regions of the feature space in which all individuals receive fixed predictions. This paradigm enables the certification of recourse for out-of-sample data, works in settings without representative datasets, and provides interpretable descriptions of individuals with fixed predictions. We develop a fast method to discover confined regions for linear classifiers using mixed-integer quadratically constrained programming. We conduct a comprehensive empirical study of confined regions across diverse applications. Our results highlight that existing pointwise verification methods fail to anticipate future individuals with fixed predictions, while our method both identifies them and provides an interpretable description.

算法公平可解释性预测验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。