arXiv:2508.06155cs.CL2025-08被引 12

提出可解释方法,精准识别大模型隐性偏见

Semantic and Structural Analysis of Implicit Biases in Large Language Models: An Interpretable Approach

  • 融合嵌套语义与上下文对比,挖掘模型输出中的隐藏偏见特征
  • 在StereoSet上实现高准确率检测,对相似文本偏见差异敏感
  • 结构透明可解释,适合对生成内容可信度要求高的场景

本文针对大语言模型生成过程中可能产生的隐性刻板印象问题,提出一种可解释的偏见检测方法,旨在识别模型输出中难以通过显式语言特征捕捉的潜在社会偏见。该方法结合嵌套语义表示与上下文对比机制,从模型输出的向量空间结构中提取潜在偏见特征。通过注意力权重扰动分析模型对特定社会属性词的敏感性,揭示偏见形成的关键语义路径。为验证方法有效性,研究使用涵盖性别、职业、宗教和种族等多个维度的StereoSet数据集,评估指标包括偏见检测准确率、语义一致性和上下文敏感性。实验结果表明,该方法在多维度上均表现出优异的检测性能,能准确区分语义相近文本间的偏见差异,同时保持高语义一致性与输出稳定性。其结构设计具有高度可解释性,有助于揭示语言模型内部的偏见关联机制,为偏见检测提供更透明可靠的技 术基础,适用于对生成内容可信度要求较高的实际应用场景。

原文摘要 · Abstract (English)

This paper addresses the issue of implicit stereotypes that may arise during the generation process of large language models. It proposes an interpretable bias detection method aimed at identifying hidden social biases in model outputs, especially those semantic tendencies that are not easily captured through explicit linguistic features. The method combines nested semantic representation with a contextual contrast mechanism. It extracts latent bias features from the vector space structure of model outputs. Using attention weight perturbation, it analyzes the model's sensitivity to specific social attribute terms, thereby revealing the semantic pathways through which bias is formed. To validate the effectiveness of the method, this study uses the StereoSet dataset, which covers multiple stereotype dimensions including gender, profession, religion, and race. The evaluation focuses on several key metrics, such as bias detection accuracy, semantic consistency, and contextual sensitivity. Experimental results show that the proposed method achieves strong detection performance across various dimensions. It can accurately identify bias differences between semantically similar texts while maintaining high semantic alignment and output stability. The method also demonstrates high interpretability in its structural design. It helps uncover the internal bias association mechanisms within language models. This provides a more transparent and reliable technical foundation for bias detection. The approach is suitable for real-world applications where high trustworthiness of generated content is required.

偏见检测可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。