arXiv:2510.24488cs.CLcs.AI2025-10被引 2

用词关联网络测大模型隐性偏见,可与人类对比

A word association network methodology for evaluating implicit biases in LLMs compared to humans

  • 通过模拟语义启动构建词关联网络,挖掘模型隐性关系
  • 发现大模型与人类在性别、宗教等偏见上既相似又不同
  • 适合关注模型公平性与认知对齐的研究者使用

随着大语言模型(LLMs)日益融入日常生活,其内在社会偏见仍是重大关切。由于这些偏见常为隐性而非显性,开发评估模型隐性知识表征的方法至关重要。本文提出一种基于语义启动模拟的词关联网络方法,通过提示工程挖掘大模型中编码的隐性关系结构,实现对偏见的量化与定性评估。该方法突破传统提示法局限,可直接比较多个大模型与人类在性别、宗教、种族、性取向及政党等社会议题上的偏见表现。实验结果揭示模型与人类偏见存在共现与差异,为理解模型认知对齐提供新视角。本方法构建了系统化、可扩展、通用的偏见评估框架,助力实现透明且负责任的语言技术。

原文摘要 · Abstract (English)

As Large language models (LLMs) become increasingly integrated into our lives, their inherent social biases remain a pressing concern. Detecting and evaluating these biases can be challenging because they are often implicit rather than explicit in nature, so developing evaluation methods that assess the implicit knowledge representations of LLMs is essential. We present a novel word association network methodology for evaluating implicit biases in LLMs based on simulating semantic priming within LLM-generated word association networks. Our prompt-based approach taps into the implicit relational structures encoded in LLMs, providing both quantitative and qualitative assessments of bias. Unlike most prompt-based evaluation methods, our method enables direct comparisons between various LLMs and humans, providing a valuable point of reference and offering new insights into the alignment of LLMs with human cognition. To demonstrate the utility of our methodology, we apply it to both humans and several widely used LLMs to investigate social biases related to gender, religion, ethnicity, sexual orientation, and political party. Our results reveal both convergences and divergences between LLM and human biases, providing new perspectives on the potential risks of using LLMs. Our methodology contributes to a systematic, scalable, and generalizable framework for evaluating and comparing biases across multiple LLMs and humans, advancing the goal of transparent and socially responsible language technologies.

偏见检测大模型评估认知对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。