提出多维度公平性评估框架,揭示毒性检测中的隐性偏差问题。
Fair and Calibrated Toxicity Detection with Robust Training and Abstention

- 构建排名、校准与拒答三轴评估体系,统一衡量模型公平性
- 发现ERM虽整体校准良好,但子群体间校准误差达0.029至0.134
- 拒答机制本身存在不公平,对身份提及内容保护不足
毒性分类的公平性涉及排名、校准和拒答三个相互关联的维度。训练阶段干预与事后安全机制无法独立评估,前者决定了后者的效果。我们比较了经验风险最小化(ERM)、实例级重加权和组别DRO在这些维度上的表现,并结合温度缩放、基于置信度的拒答以及按身份优化阈值的方法。评估采用子组AUC、BPSN/BNSP AUC、误差差距和带自举置信区间(n=1000)的各子组期望校准误差(ECE)。研究发现:(1)校准差异是一种隐蔽的不公平,ERM虽整体校准极佳(ECE=0.013),但在所有身份子组中显著失准(+0.029至+0.134);(2)训练干预仅重塑而非消除偏差,重加权ERM提升排名性能(BPSN AUC提升0.06至0.12),但使校准-公平差距扩大至+0.232;组别DRO消除校准差异,但导致全局均匀失准(ECE=0.118);(3)事后方法继承训练缺陷,温度缩放因非均匀失准而失效,基于置信度的拒答在ERM下有效,但在DRO下失效,拒答率随退避增加而上升;(4)拒答机制本身不公平,对背景内容保护远优于提及身份的内容。我们认为SRAI公平性需多轴框架,仅在聚合排名上差异的方法,其失败模式可能导致真实世界伤害的显著差异。
原文摘要 · Abstract (English)
Fairness in toxicity classification involves three integrated axes: ranking, calibration, and abstention. Training-time interventions and post-hoc safety mechanisms cannot be evaluated independently because the former determines the efficacy of the latter. We compare Empirical Risk Minimization (ERM), instance-level reweighting, and Group DRO across these axes, combined with temperature scaling, confidence-based abstention, and per-identity threshold optimization. Evaluation uses subgroup AUC, BPSN/BNSP AUC, error gaps, and per-subgroup Expected Calibration Error (ECE) with bootstrap CIs ($n = 1000$). We report four findings. (1) Calibration disparity is a hidden fairness violation. ERM has near-perfect aggregate calibration ($0.013$) but is significantly miscalibrated across all identity subgroups ($+0.029$ to $+0.134$). (2) Training interventions reshape rather than eliminate disparity. Reweighted ERM improves ranking (BPSN AUC $+0.06$ to $+0.12$) but worsens the calibration-fairness gap by up to $+0.232$. Group DRO eliminates calibration disparity but only by becoming uniformly miscalibrated globally (ECE $0.118$). (3) Post-hoc methods inherit training failure modes. Temperature scaling fails because miscalibration is non-uniform. Confidence-based abstention works under ERM but breaks under DRO, where the risk-coverage curve rises with deferral. (4) Abstention itself is unfair. Confidence-based deferral helps background content far more than identity-mentioning content. We argue that SRAI fairness requires a multi-axis framework: methods that differ only in aggregate ranking can differ sharply in failure modes that determine real-world harm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。