arXiv:2412.04914cs.LGstat.ML2024-12被引 5

通过独立性约束提升流程预测中的群体公平性

Achieving Group Fairness through Independence in Predictive Process Monitoring

  • 用分布距离度量确保预测与敏感群体无关
  • 新损失函数平衡预测准确率与公平性
  • 适用于需避免偏见的工业流程决策场景

预测性流程监控旨在预测正在进行的流程实例未来状态,如案件结果。近年来,机器学习模型在该领域的应用受到广泛关注。但历史执行数据可能包含偏见或不公平行为,这些偏见会被编码到训练模型中。当此类模型用于新案例的决策或干预时,可能延续不良行为。本文通过研究独立性来解决预测性流程监控中的群体公平性问题,即确保预测结果不受敏感群体成员身份影响。采用人口统计均等性指标ΔDP,以及近期提出的、不依赖阈值的基于分布的替代度量。此外,提出一种由二元交叉熵和基于分布的损失(Wasserstein)组成的复合损失函数,以训练在预测性能与公平性之间可定制权衡的模型。通过受控实验验证了公平性度量与复合损失函数的有效性。

原文摘要 · Abstract (English)

Predictive process monitoring focuses on forecasting future states of ongoing process executions, such as predicting the outcome of a particular case. In recent years, the application of machine learning models in this domain has garnered significant scientific attention. When using historical execution data, which may contain biases or exhibit unfair behavior, these biases may be encoded into the trained models. Consequently, when such models are deployed to make decisions or guide interventions for new cases, they risk perpetuating this unwanted behavior. This work addresses group fairness in predictive process monitoring by investigating independence, i.e. ensuring predictions are unaffected by sensitive group membership. We explore independence through metrics for demographic parity such as $Δ$DP, as well as recently introduced, threshold-independent distribution-based alternatives. Additionally, we propose a composite loss function existing of binary cross-entropy and a distribution-based loss (Wasserstein) to train models that balance predictive performance and fairness, and allow for customizable trade-offs. The effectiveness of both the fairness metrics and the composite loss functions is validated through a controlled experimental setup.

流程监控公平性预测模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。