研究置信预测中公平性问题,发现均衡预测集大小更利于公平决策。
Beyond Procedure: Substantive Fairness in Conformal Prediction
- 提出从决策全流程评估公平性,超越传统程序公平
- 理论证明预测集大小差异可分解为可解释成分
- 用大模型辅助评估公平性,适合关注公平的开发者
置信预测(Conformal Prediction, CP)为机器学习模型提供无需分布假设的不确定性量化,但其在下游决策中的公平性影响尚未深入探讨。本文超越将CP视为独立操作的视角(程序公平),从完整决策流程出发,分析结果公平性(substantive fairness)——即下游结果的公平程度。理论上,我们推导出预测集大小差异的上界,将其分解为可解释成分,揭示标签聚类的CP如何控制方法本身导致的不公平。为支持大规模实证分析,我们引入大模型辅助评估器,跨多种模态近似人类对公平性的判断。实验表明,标签聚类的CP通常在效用与公平性之间取得良好平衡,且符合理论预期地降低预测集大小差异。最后,我们发现预测集大小均等比覆盖率更强烈相关于结果公平性提升,为设计更公平的置信预测系统提供了实践指导。代码已开源。
原文摘要 · Abstract (English)
Conformal prediction (CP) offers distribution-free uncertainty quantification for machine learning models, yet its interplay with fairness in downstream decision-making remains underexplored. Moving beyond CP as a standalone operation (procedural fairness), we analyze the holistic decision-making pipeline to evaluate substantive fairness-the equity of downstream outcomes. Theoretically, we derive an upper bound that decomposes prediction-set size disparity into interpretable components, clarifying how label-clustered CP helps control method-driven contributions to unfairness. To facilitate scalable empirical analysis, we introduce an LLM-in-the-loop evaluator that approximates human assessment of substantive fairness across diverse modalities. Our experiments show that label-clustered CP often provides a favorable balance between utility and substantive fairness, while reducing set-size disparities in line with our theory. Finally, we empirically show that equalized set sizes, rather than coverage, strongly correlate with improved substantive fairness, enabling practitioners to design more fair CP systems. Our code is available at https://github.com/layer6ai-labs/llm-in-the-loop-conformal-fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。