arXiv:2503.05047cs.CL2025-03被引 1

用大模型生成的文本数据含偏见和标注缺陷,需警惕。

Biases in Large Language Model-Elicited Text: A Case Study in Natural Language Inference

  • 用GPT-4等大模型重制自然语言推理数据集
  • 大模型生成数据中性别、种族、年龄相关偏见明显
  • 适合关注大模型数据质量与社会偏见的研究者

我们测试了由大语言模型(LLMs)生成的NLP数据集是否包含标注缺陷和社交偏见,类似于众包标注的数据集。我们使用GPT-4、Llama-2 70b Chat和Mistral 7b Instruct重制了斯坦福自然语言推理语料库的一部分。通过训练仅基于假设的分类器,判断大模型生成的NLI数据集是否存在标注偏差。随后,利用点互信息(PMI)识别每个数据集中与性别、种族和年龄相关术语关联的词汇。在大模型生成的NLI数据集上,微调后的BERT假设分类器准确率达86%-96%。分析进一步揭示了大模型生成数据中的标注缺陷与刻板偏见。

原文摘要 · Abstract (English)

We test whether NLP datasets created with Large Language Models (LLMs) contain annotation artifacts and social biases like NLP datasets elicited from crowd-source workers. We recreate a portion of the Stanford Natural Language Inference corpus using GPT-4, Llama-2 70b for Chat, and Mistral 7b Instruct. We train hypothesis-only classifiers to determine whether LLM-elicited NLI datasets contain annotation artifacts. Next, we use pointwise mutual information to identify the words in each dataset that are associated with gender, race, and age-related terms. On our LLM-generated NLI datasets, fine-tuned BERT hypothesis-only classifiers achieve between 86-96% accuracy. Our analyses further characterize the annotation artifacts and stereotypical biases in LLM-generated datasets.

大模型偏见自然语言推理数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。