arXiv:2601.06761cs.SEcs.LG2026-01

用对比判断数据评估模型公平性,降低人工标注负担。

Comparative Separation: Evaluating Separation on Comparative Judgment Test Data

  • 提出对比分离概念,基于成对比较数据评估公平性。
  • 理论证明在二分类中,对比分离等价于传统分离标准。
  • 实证显示对比判断可减少数据量需求,提升评估效率。

本研究旨在为软件工程领域提供新方法,提出对比分离这一新型群体公平性概念,用于在对比判断测试数据上评估机器学习软件的公平性。随着机器学习广泛应用于高风险决策场景,确保模型对不同敏感群体表现一致(满足分离准则)成为开发者的责任。然而,传统分离评估需每个样本的真值标签,成本高昂。本文探索在仅提供成对比较判断(如A优于B)的数据上是否可行。根据比较判断定律,成对比较比评分或分类标注更轻负担。本文首次定义了对比分离及其评估指标,并在理论上与实验中证明:在二分类任务中,对比分离与分离等价。进一步分析了达到相同统计效力所需的样本点数与数据对数,结果表明对比判断策略在资源利用上更具优势。这是首个探讨对比判断数据下公平性评估的研究,证实了其可行性与实际价值。

原文摘要 · Abstract (English)

This research seeks to benefit the software engineering society by proposing comparative separation, a novel group fairness notion to evaluate the fairness of machine learning software on comparative judgment test data. Fairness issues have attracted increasing attention since machine learning software is increasingly used for high-stakes and high-risk decisions. It is the responsibility of all software developers to make their software accountable by ensuring that the machine learning software do not perform differently on different sensitive groups -- satisfying the separation criterion. However, evaluation of separation requires ground truth labels for each test data point. This motivates our work on analyzing whether separation can be evaluated on comparative judgment test data. Instead of asking humans to provide the ratings or categorical labels on each test data point, comparative judgments are made between pairs of data points such as A is better than B. According to the law of comparative judgment, providing such comparative judgments yields a lower cognitive burden for humans than providing ratings or categorical labels. This work first defines the novel fairness notion comparative separation on comparative judgment test data, and the metrics to evaluate comparative separation. Then, both theoretically and empirically, we show that in binary classification problems, comparative separation is equivalent to separation. Lastly, we analyze the number of test data points and test data pairs required to achieve the same level of statistical power in the evaluation of separation and comparative separation, respectively. This work is the first to explore fairness evaluation on comparative judgment test data. It shows the feasibility and the practical benefits of using comparative judgment test data for model evaluations.

公平性评估对比判断机器学习软件工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。