arXiv:2504.14212cs.CL2025-04EMNLP被引 2

通过检测保护属性与情感倾向,分析并缓解大模型预训练数据中的社会偏见。

Bias Analysis and Mitigation through Protected Attribute Detection and Regard Classification

  • 先识别文本中的受保护属性(如性别、种族),再判断对其态度
  • 在Common Crawl数据上发现显著语言极性偏见,尤其对少数群体
  • 适用于研究模型偏见或开发公平性工具的研究者

大规模语言模型(LLMs)通过海量文本预训练获得通用语言知识,但这些预训练数据主要来自网络爬取文本,常包含不良社会偏见,可能被模型继承甚至放大。本研究提出一种高效有效的标注流程,用于分析预训练语料库中的社会偏见。该流程包括受保护属性检测以识别多样化人口特征,以及态度分类以分析对各类属性的语言极性。实验聚焦于最具代表性的预训练数据集Common Crawl,验证了所提偏见分析与缓解方法的有效性。

原文摘要 · Abstract (English)

Large language models (LLMs) acquire general linguistic knowledge from massive-scale pretraining. However, pretraining data mainly comprised of web-crawled texts contain undesirable social biases which can be perpetuated or even amplified by LLMs. In this study, we propose an efficient yet effective annotation pipeline to investigate social biases in the pretraining corpora. Our pipeline consists of protected attribute detection to identify diverse demographics, followed by regard classification to analyze the language polarity towards each attribute. Through our experiments, we demonstrate the effect of our bias analysis and mitigation measures, focusing on Common Crawl as the most representative pretraining corpus.

模型偏见公平性自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。