arXiv:2501.16549cs.CYcs.LG2025-01被引 3

解决模型预测不一致问题,让不同模型达成共识。

Reconciling Predictive Multiplicity in Practice

  • 基于模型分歧反向修正错误预测,提升可靠性。
  • 在5个公平性数据集上验证算法有效,降低预测冲突。
  • 首次拓展至因果推断,适用于处理治疗效应分歧。

许多机器学习应用需预测个体概率(如患病风险),但真实概率未知,导致同一数据集训练的不同模型对某些个体预测不一致,即模型多重性(MM)现象。Roth等提出的Reconcile算法通过利用模型间的分歧来识别并改进至少一个模型。本文在COMPAS、Communities and Crime、Adult、Statlog(German Credit Data)和ACS数据集上实证分析该算法,评估其在模型多重性研究中的定位,并与现有方案对比,验证其有效性。同时探讨算法的理论与实践优化方向。进一步将算法扩展至因果推断场景,因不同估计器对特定平均处理效应(CATE)值也可能存在分歧。本文首次实现Reconcile算法在因果推断中的应用,分析其理论性质并开展实证测试,结果证实其在多种场景下的实用性。

原文摘要 · Abstract (English)

Many machine learning applications predict individual probabilities, such as the likelihood that a person develops a particular illness. Since these probabilities are unknown, a key question is how to address situations in which different models trained on the same dataset produce varying predictions for certain individuals. This issue is exemplified by the model multiplicity (MM) phenomenon, where a set of comparable models yield inconsistent predictions. Roth, Tolbert, and Weinstein recently introduced a reconciliation procedure, the Reconcile algorithm, to address this problem. Given two disagreeing models, the algorithm leverages their disagreement to falsify and improve at least one of the models. In this paper, we empirically analyze the Reconcile algorithm using five widely-used fairness datasets: COMPAS, Communities and Crime, Adult, Statlog (German Credit Data), and the ACS Dataset. We examine how Reconcile fits within the model multiplicity literature and compare it to existing MM solutions, demonstrating its effectiveness. We also discuss potential improvements to the Reconcile algorithm theoretically and practically. Finally, we extend the Reconcile algorithm to the setting of causal inference, given that different competing estimators can again disagree on specific causal average treatment effect (CATE) values. We present the first extension of the Reconcile algorithm in causal inference, analyze its theoretical properties, and conduct empirical tests. Our results confirm the practical effectiveness of Reconcile and its applicability across various domains.

模型多重性因果推断公平性算法修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。