arXiv:2508.01412cs.CL2025-08AAAI被引 4

发现大模型生成中隐藏的社会偏见关联,助力识别未被察觉的歧视性表达。

Bias Association Discovery Framework for Open-Ended LLM Generations

  • 基于开放生成文本,自动挖掘身份与概念间的潜在偏见关联。
  • 在多模型、多场景实验中揭示了多种未被预设的偏见模式。
  • 适合研究者与开发者用于系统检测和分析大模型中的隐性偏见。

大型语言模型(LLMs)中嵌入的社会偏见引发严重关切,导致表征伤害——即对特定人口群体的不公平或扭曲表述,这些偏见可能通过生成语言以微妙方式体现。现有评估方法通常依赖预定义的身份-概念关联,难以发现新的或意外的偏见形式。本文提出偏见关联发现框架(BADF),一种系统化方法,可从开放生成的LLM输出中提取已知及此前未识别的身份-概念关联。通过覆盖多个模型与多样真实场景的综合实验,BADF实现了对刻画人口身份的多样化概念的稳健映射与分析。研究结果深化了对开放生成中偏见的理解,并提供了一种可扩展的工具,用于识别与分析大模型中的偏见关联。

原文摘要 · Abstract (English)

Social biases embedded in Large Language Models (LLMs) raise critical concerns, resulting in representational harms -- unfair or distorted portrayals of demographic groups -- that may be expressed in subtle ways through generated language. Existing evaluation methods often depend on predefined identity-concept associations, limiting their ability to surface new or unexpected forms of bias. In this work, we present the Bias Association Discovery Framework (BADF), a systematic approach for extracting both known and previously unrecognized associations between demographic identities and descriptive concepts from open-ended LLM outputs. Through comprehensive experiments spanning multiple models and diverse real-world contexts, BADF enables robust mapping and analysis of the varied concepts that characterize demographic identities. Our findings advance the understanding of biases in open-ended generation and provide a scalable tool for identifying and analyzing bias associations in LLMs.

偏见检测大模型语言模型社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。