arXiv:2502.20898cs.CL2025-02

构建数据库评估GPT-4o性别偏见,推动公平性研究

A database to support the evaluation of gender biases in GPT-4o output

  • 设计可复现的数据库构建方法,支持多维度评估性别偏见
  • 不仅能测偏见程度,还能分析其生成机制与表现形式
  • 适合关注AI伦理、公平性评测的研究者与开发者

大型语言模型(LLMs)的广泛应用带来了用户与社会的伦理风险。其中,生成强化或加剧对弱势群体伤害的不公平语言输出,是显著的伦理问题,尤其体现在性别偏见方面(Weidinger et al., 2022; Bender et al., 2021; Kotek et al., 2023)。因此,评估LLM输出在性别方面的公平性已成为研究热点。为推进该领域研究,促进对规范基础与评估方法的讨论,并提升相关研究的可复现性,本文提出一种新颖的数据库构建方法,使对大模型生成语言中性别相关偏见的评估超越单纯测量中立化程度,实现更全面、深入的分析。

原文摘要 · Abstract (English)

The widespread application of Large Language Models (LLMs) involves ethical risks for users and societies. A prominent ethical risk of LLMs is the generation of unfair language output that reinforces or exacerbates harm for members of disadvantaged social groups through gender biases (Weidinger et al., 2022; Bender et al., 2021; Kotek et al., 2023). Hence, the evaluation of the fairness of LLM outputs with respect to such biases is a topic of rising interest. To advance research in this field, promote discourse on suitable normative bases and evaluation methodologies, and enhance the reproducibility of related studies, we propose a novel approach to database construction. This approach enables the assessment of gender-related biases in LLM-generated language beyond merely evaluating their degree of neutralization.

性别偏见LLM评估数据库构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。