arXiv:2505.17533cs.LGcs.AI2025-05

通过可解释的表征差异,让算法帮人减少决策偏差。

Learning Representational Disparities

  • 用神经网络建模人类决策中实际与理想表征的差异
  • 在德国信贷等3个数据集上显著降低结果不平等
  • 结果可解释,适合用于设计针对性干预措施

我们提出一种公平机器学习算法,用于建模观察到的人类决策与理想决策之间的可解释差异,后者旨在减少受人类决策影响的下游结果中的不平等。以往工作在学习公平表征时未考虑决策过程中的结果影响。我们假设结果不平等源于观察者与理想决策者对输入的不同表征,称为表征差异。目标是学习可解释的表征差异,可通过特定提示修正人类决策,从而缓解下游结果的不平等;这被建模为多目标优化问题。在合理简化假设下,我们证明该神经网络模型所学的权重可完全消除结果不平等。我们在德国信贷、成人和遗产健康三个真实数据集上验证了目标有效性与结果可解释性。

原文摘要 · Abstract (English)

We propose a fair machine learning algorithm to model interpretable differences between observed and desired human decision-making, with the latter aimed at reducing disparity in a downstream outcome impacted by the human decision. Prior work learns fair representations without considering the outcome in the decision-making process. We model the outcome disparities as arising due to the different representations of the input seen by the observed and desired decision-maker, which we term representational disparities. Our goal is to learn interpretable representational disparities which could potentially be corrected by specific nudges to the human decision, mitigating disparities in the downstream outcome; we frame this as a multi-objective optimization problem using a neural network. Under reasonable simplifying assumptions, we prove that our neural network model of the representational disparity learns interpretable weights that fully mitigate the outcome disparity. We validate objectives and interpret results using real-world German Credit, Adult, and Heritage Health datasets.

公平学习可解释性表征差异决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。