arXiv:2511.21711cs.CLcs.CY2025-11被引 1

检测并缓解大模型中的性别与种族偏见,提升生成内容的公平性。

Addressing Stereotypes in Large Language Models: A Critical Examination and Mitigation

  • 采用多维度评估框架,结合关键词与上下文分析偏见。
  • 微调后模型在隐性偏见测试中性能提升最高达20%。
  • 适合关注AI伦理、模型公平性的研究人员和开发者。

大型语言模型(如ChatGPT)因自然语言处理技术进步而广泛应用,但其训练数据中的显性和隐性偏见可能引发社会、文化、宗教等层面的刻板印象。本研究通过StereoSet和CrowSPairs等专用基准,评估BERT、GPT 3.5、ADA等多种生成模型的偏见表现。采用三重分析方法全面识别偏见来源与程度。结果显示,微调模型在性别偏见上仍存不足,但在种族偏见识别与规避方面表现较好;模型常过度依赖提示词,缺乏对输出真实性的理解能力。为提升性能,引入基于微调、不同提示策略及基准数据增强的强化学习方法,结果表明模型在跨数据集测试中具备良好适应性,隐性偏见基准性能最高提升20%。

原文摘要 · Abstract (English)

Large Language models (LLMs), such as ChatGPT, have gained popularity in recent years with the advancement of Natural Language Processing (NLP), with use cases spanning many disciplines and daily lives as well. LLMs inherit explicit and implicit biases from the datasets they were trained on; these biases can include social, ethical, cultural, religious, and other prejudices and stereotypes. It is important to comprehensively examine such shortcomings by identifying the existence and extent of such biases, recognizing the origin, and attempting to mitigate such biased outputs to ensure fair outputs to reduce harmful stereotypes and misinformation. This study inspects and highlights the need to address biases in LLMs amid growing generative Artificial Intelligence (AI). We utilize bias-specific benchmarks such StereoSet and CrowSPairs to evaluate the existence of various biases in many different generative models such as BERT, GPT 3.5, and ADA. To detect both explicit and implicit biases, we adopt a three-pronged approach for thorough and inclusive analysis. Results indicate fine-tuned models struggle with gender biases but excel at identifying and avoiding racial biases. Our findings also illustrated that despite some cases of success, LLMs often over-rely on keywords in prompts and its outputs. This demonstrates the incapability of LLMs to attempt to truly understand the accuracy and authenticity of its outputs. Finally, in an attempt to bolster model performance, we applied an enhancement learning strategy involving fine-tuning, models using different prompting techniques, and data augmentation of the bias benchmarks. We found fine-tuned models to exhibit promising adaptability during cross-dataset testing and significantly enhanced performance on implicit bias benchmarks, with performance gains of up to 20%.

大模型偏见公平性提示工程数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。