arXiv:2606.10632cs.LGcs.AI2026-06KDD

提出固定阈值评估方法,让多任务学习公平性比较更可靠。

Is Fairness Truly Fair? Towards Reliable Lipschitz Fairness in Multi-Task Learning via Fixed-\texorpdfstring{$δ$}{delta} Alignment

论文配图:Is Fairness Truly Fair? Towards Reliable Lipschitz Fairness in Multi-Task Learning via Fixed-\texorpdfstring{$δ$}{delta} Alignment
图 1 · 摘自论文原文
  • 用统一阈值替代各模型自定阈值,避免评估偏差。
  • 在临床时序和深度预测任务中验证了公平性-性能权衡。
  • 适合关注多任务学习公平性评估的算法研究者。

Lipschitz风格的个体公平性要求语义相似样本获得相似预测,但在多任务学习中,其评估常受方法导致的表征尺度影响。本文揭示了阈值混淆问题:当审计容差基于各模型自身表征距离确定时,不同算法在不同语义阈值下被比较。阈值漂移分析表明偏见排名可能变化,并识别出排名保持的充分条件。我们提出ReLiF框架,将评估时的固定δ审计与训练时的受控正则化分离。ReLiF采用共享参考容差实现可比审计,并使用违规率反馈控制器维持Lipschitz代理活跃,防止其主导随机训练。该工作还提供了关于阈值漂移、参考容差选择及Huber化训练代理与其非光滑正边距版本关系的分析。在临床时间序列基准和NYUv2密集预测任务上的实验表明,固定δ审计揭示了方法依赖阈值会掩盖的效用-公平性权衡。在以ResNet50为骨干的NYUv2上,ReLiF在保持竞争力效用的同时显著降低对齐偏见;在临床基准上,ReLiF实现可控的公平性正则化权衡,而固定δ审计显示任务平衡基线有时能取得更低偏见,且真实的效用-公平性权衡依然存在。这些结果支持固定δ审计作为多任务学习中Lipschitz公平性评估的语义一致协议。

原文摘要 · Abstract (English)

Lipschitz-style individual fairness formalizes the idea that semantically similar examples should receive similar predictions, but its evaluation in multi-task learning (MTL) can be confounded by method-induced representation scales. This paper identifies threshold confounding: when the auditing tolerance is derived from each model's own representation distances, different algorithms are compared under different semantic thresholds. A threshold-drift analysis further shows how Bias rankings can change and identifies sufficient conditions for ranking preservation. We propose \textbf{ReLiF}, a reliability-aware framework that separates evaluation-time fixed-$δ$ auditing from training-time controlled regularization. ReLiF uses a shared reference tolerance for comparable auditing and a violation-rate feedback controller to keep the Lipschitz surrogate active without letting it dominate stochastic training. This work also develops supporting analysis for threshold drift, reference-tolerance selection, and the relationship between the huberized training surrogate and its unsmoothed positive-margin counterpart. Experiments on clinical time-series benchmarks and NYUv2 (NYU Depth V2) dense prediction show that fixed-$δ$ auditing exposes utility--fairness trade-offs that method-dependent thresholds can obscure. On NYUv2 with a ResNet50 backbone, ReLiF achieves competitive utility while substantially reducing aligned bias under shared fixed thresholds. On clinical benchmarks, ReLiF yields controlled fairness-regularized trade-offs, while fixed-$δ$ auditing reveals that task-balancing baselines can sometimes achieve lower bias and that genuine utility--fairness trade-offs persist. These results support fixed-$δ$ auditing as a semantically consistent protocol for evaluating Lipschitz fairness in MTL.

公平性多任务学习评估方法深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。