arXiv:2503.23243cs.CLcs.AI2025-03被引 8

评估大模型在敏感议题标注中是否代表多元观点

Evaluating how LLM annotations represent diverse views on contentious topics

  • 对比多个大模型在四个数据集上的标注表现
  • 模型偏差方向一致,且标注难度比人口属性更影响一致性
  • 提醒研究者需结合任务难度做公平性评估

研究人员提出使用生成式大语言模型(LLMs)进行数据标注。现有文献强调这类模型在性能上优于其他自然语言模型,甚至在多个指标上超越人类。尽管已有研究关注各类应用中的偏见,但对生成式LLMs在主观标注任务中的偏见关注较少。这种偏见可能导致模型标签偏向多数群体,忽视多元观点。本文在四个数据集的四项标注任务中评估了LLMs对多元观点的代表性。结果显示,不同模型在相同数据集的同一个人口类别上表现出一致的偏差方向,而人类标注者之间的分歧(即任务难度)更能预测模型与人类的一致性。结论指出:研究者和实践者在使用LLMs进行自动标注时,必须进行情境化公平性评估,仅靠模型选择无法解决偏见问题,且应将任务难度纳入偏见分析。

原文摘要 · Abstract (English)

Researchers have proposed the use of generative large language models (LLMs) to label data for research and applied settings. This literature emphasizes the improved performance of these models relative to other natural language models, noting that generative LLMs typically outperform other models and even humans across several metrics. Previous literature has examined bias across many applications and contexts, but less work has focused specifically on bias in generative LLMs' responses to subjective annotation tasks. This bias could result in labels applied by LLMs that disproportionately align with majority groups over a more diverse set of viewpoints. In this paper, we evaluate how LLMs represent diverse viewpoints on these contentious tasks. Across four annotation tasks on four datasets, we show that LLMs do not show systematic substantial disagreement with annotators on the basis of demographics. Rather, we find that multiple LLMs tend to be biased in the same directions on the same demographic categories within the same datasets. Moreover, the disagreement between human annotators on the labeling task -- a measure of item difficulty -- is far more predictive of LLM agreement with human annotators. We conclude with a discussion of the implications for researchers and practitioners using LLMs for automated data annotation tasks. Specifically, we emphasize that fairness evaluations must be contextual, model choice alone will not solve potential issues of bias, and item difficulty must be integrated into bias assessments.

大模型偏见自动标注公平性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。