arXiv:2509.22699cs.CL2025-09EMNLP被引 2

用置信度衡量模型对弱势群体内容的偏见,发现高准确率下信心低

Are you sure? Measuring models bias in content moderation through uncertainty

  • 用校准预测技术计算模型对不同群体标注的不确定性
  • 11个模型对少数群体内容准确率高但置信度低,显示隐性偏见
  • 适合关注模型公平性与安全部署的研究者使用

自动内容审核对社交媒体安全至关重要。基于语言模型的分类器被广泛采用,但已被证明会延续种族与社会偏见。尽管已有多个资源和基准数据集用于应对该问题,衡量模型在内容审核中的公平性仍是开放难题。本文提出一种无监督方法,基于模型对来自弱势群体标注消息的不确定性进行评测。我们利用校准预测技术计算的不确定性作为代理指标,分析11个模型对女性和非白人标注者数据的偏差,并观察其与性能指标(如F1分数)的差异。结果显示,某些预训练模型虽能高精度预测少数群体标注的内容,但其预测置信度却较低。因此,通过测量模型置信度,可识别哪些群体在预训练模型中代表性不足,从而在实际应用前推动模型去偏过程。

原文摘要 · Abstract (English)

Automatic content moderation is crucial to ensuring safety in social media. Language Model-based classifiers are being increasingly adopted for this task, but it has been shown that they perpetuate racial and social biases. Even if several resources and benchmark corpora have been developed to challenge this issue, measuring the fairness of models in content moderation remains an open issue. In this work, we present an unsupervised approach that benchmarks models on the basis of their uncertainty in classifying messages annotated by people belonging to vulnerable groups. We use uncertainty, computed by means of the conformal prediction technique, as a proxy to analyze the bias of 11 models against women and non-white annotators and observe to what extent it diverges from metrics based on performance, such as the $F_1$ score. The results show that some pre-trained models predict with high accuracy the labels coming from minority groups, even if the confidence in their prediction is low. Therefore, by measuring the confidence of models, we are able to see which groups of annotators are better represented in pre-trained models and lead the debiasing process of these models before their effective use.

内容审核模型偏见置信度公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。