用心理学量表分析大模型中的性别偏见,直观展示其隐藏歧视。
Profiling Bias in LLMs: Stereotype Dimensions in Contextual Word Embeddings
- 基于社会心理学量表构建偏见评估维度
- 发现12个大模型在不同上下文层中存在性别刻板印象
- 适合研究人员和公众理解模型偏见的可视化工具
大型语言模型(LLMs)是当前人工智能成功的基础,但不可避免地存在偏见。为有效传达风险并推动缓解措施,需要对模型的歧视性特征进行清晰、直观的描述,以适应各类AI受众。本文基于社会心理学研究中的词典,提出以刻板印象维度为基础的偏见画像。我们沿着这些维度,在不同上下文和模型层中分析了性别偏见,并为12个不同的大模型生成了刻板印象画像,证明了其直观性和在暴露与可视化偏见方面的实用价值。
原文摘要 · Abstract (English)
Large language models (LLMs) are the foundation of the current successes of artificial intelligence (AI), however, they are unavoidably biased. To effectively communicate the risks and encourage mitigation efforts these models need adequate and intuitive descriptions of their discriminatory properties, appropriate for all audiences of AI. We suggest bias profiles with respect to stereotype dimensions based on dictionaries from social psychology research. Along these dimensions we investigate gender bias in contextual embeddings, across contexts and layers, and generate stereotype profiles for twelve different LLMs, demonstrating their intuition and use case for exposing and visualizing bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。