arXiv:2508.14576cs.LG2025-08AAAI

不同密度比估计方法让回归公平性测量结果差异巨大,可能误导判断。

A Comprehensive Evaluation of the Sensitivity of Density-Ratio Estimation Based Fairness Measurement in Regression

  • 用多种密度比估计方法构建公平性测量框架
  • 实验发现核心算法选择影响公平性评估结果,甚至导致矛盾结论
  • 提醒研究者谨慎选择评估工具,适合关注算法偏见的从业者

机器学习中的算法偏见问题日益受到关注,公平性测量成为研究热点。针对回归任务的公平性评估,近期研究将其建模为密度比估计问题,并采用基于逻辑回归的概率分类器求解。然而,密度比估计存在多种方法,现有工作未考察其对公平性测量结果的敏感性。本文构建了基于不同密度比估计核心的公平性测量方法,系统评估了不同核心对测量结果的影响。实验表明,核心算法的选择显著影响公平性评估结果,甚至导致对不同算法相对公平性的判断不一致。这一发现揭示了当前基于密度比估计的回归公平性测量存在可靠性问题,亟需进一步研究以提升其稳健性。

原文摘要 · Abstract (English)

The prevalence of algorithmic bias in Machine Learning (ML)-driven approaches has inspired growing research on measuring and mitigating bias in the ML domain. Accordingly, prior research studied how to measure fairness in regression which is a complex problem. In particular, recent research proposed to formulate it as a density-ratio estimation problem and relied on a Logistic Regression-driven probabilistic classifier-based approach to solve it. However, there are several other methods to estimate a density ratio, and to the best of our knowledge, prior work did not study the sensitivity of such fairness measurement methods to the choice of underlying density ratio estimation algorithm. To fill this gap, this paper develops a set of fairness measurement methods with various density-ratio estimation cores and thoroughly investigates how different cores would affect the achieved level of fairness. Our experimental results show that the choice of density-ratio estimation core could significantly affect the outcome of fairness measurement method, and even, generate inconsistent results with respect to the relative fairness of various algorithms. These observations suggest major issues with density-ratio estimation based fairness measurement in regression and a need for further research to enhance their reliability.

公平性测量密度比估计回归模型算法偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。