arXiv:2411.08977cs.CYcs.CL2024-11ACL被引 14

研究大模型对不同人群攻击性语言的判断偏差,发现敏感度和一致性影响更大。

Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness

  • 在五大数据集上测试模型与人类标注的一致性,共22万条标注
  • 模型对种族的判断偏差不一致,但标注者敏感度越高一致性越强
  • 文档难度高时模型与人意见更不一致,需考虑多重混杂因素

大型语言模型(LLMs)存在人口统计学偏见,但很少有研究系统评估多个数据集中的此类偏见或考虑混杂因素。本文在五个攻击性语言数据集上考察了LLM与人类标注的一致性,共包含约22万条标注。结果表明,尽管人口统计特征(尤其是种族)会影响对齐程度,但这种影响在不同数据集间不一致,且常与其他因素交织。混杂因素——如文档难度、标注者敏感度及组内一致性——对对齐模式变化的解释力超过人口统计特征本身。具体而言,标注者敏感度和组内一致性越高,对齐程度越高;而文档难度越大,对齐程度越低。研究强调了多数据集分析和混杂因素意识在构建稳健的人口统计偏见评估方法中的重要性。

原文摘要 · Abstract (English)

Large language models (LLMs) are known to exhibit demographic biases, yet few studies systematically evaluate these biases across multiple datasets or account for confounding factors. In this work, we examine LLM alignment with human annotations in five offensive language datasets, comprising approximately 220K annotations. Our findings reveal that while demographic traits, particularly race, influence alignment, these effects are inconsistent across datasets and often entangled with other factors. Confounders -- such as document difficulty, annotator sensitivity, and within-group agreement -- account for more variation in alignment patterns than demographic traits alone. Specifically, alignment increases with higher annotator sensitivity and group agreement, while greater document difficulty corresponds to reduced alignment. Our results underscore the importance of multi-dataset analyses and confounder-aware methodologies in developing robust measures of demographic bias in LLMs.

大模型偏见人类对齐混杂因素攻击性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。