arXiv:2502.19160cs.CLcs.AI2025-02被引 11

用大模型检测语言中的刻板印象线索,提升评估的准确性与可解释性。

Detecting Linguistic Indicators for Stereotype Assessment with Large Language Models

  • 基于社会类别框架提取语言刻板印象特征,构建分类体系。
  • 通过提示学习让LLM识别句子中刻板印象的语言指标,得分函数可量化强度。
  • 模型表现随规模增长,70B级大模型效果最优,适合偏见分析研究者使用。

社会类别与刻板印象嵌入语言中,可能在大型语言模型(LLMs)中引入数据偏差。尽管已有防护措施,此类偏差仍常出现在模型输出中,可能导致表征伤害。现有自然语言处理方法缺乏对社会语言学基础的借鉴,往往缺乏客观性、精确性和可解释性。为此,本文提出一种新方法,用于检测并量化句子中刻板印象的语言指标。我们从社会类别与刻板印象传播(SCSC)框架中提取语言特征,构建分类方案,并利用上下文学习指导不同LLM对句子进行分析,判断其语言属性并提供细粒度评估依据。基于对各类语言指标重要性的实证评估,我们学习了一个评分函数,用于衡量刻板印象的语言信号强度。人工标注显示这些指标存在于刻板化语句中,且能解释刻板印象强度。实验表明,模型在识别类别标签相关语言指标上表现良好,但对关联行为与特征的判断仍存困难;增加提示中的少样本示例显著提升性能。模型规模越大表现越好:Llama-3.3-70B-Instruct 和 GPT-4 的效果优于 Mixtral-8x7B-Instruct、GPT-4-mini 和 Llama-3.1-8B-Instruct。

原文摘要 · Abstract (English)

Social categories and stereotypes are embedded in language and can introduce data bias into Large Language Models (LLMs). Despite safeguards, these biases often persist in model behavior, potentially leading to representational harm in outputs. While sociolinguistic research provides valuable insights into the formation of stereotypes, NLP approaches for stereotype detection rarely draw on this foundation and often lack objectivity, precision, and interpretability. To fill this gap, in this work we propose a new approach that detects and quantifies the linguistic indicators of stereotypes in a sentence. We derive linguistic indicators from the Social Category and Stereotype Communication (SCSC) framework which indicate strong social category formulation and stereotyping in language, and use them to build a categorization scheme. To automate this approach, we instruct different LLMs using in-context learning to apply the approach to a sentence, where the LLM examines the linguistic properties and provides a basis for a fine-grained assessment. Based on an empirical evaluation of the importance of different linguistic indicators, we learn a scoring function that measures the linguistic indicators of a stereotype. Our annotations of stereotyped sentences show that these indicators are present in these sentences and explain the strength of a stereotype. In terms of model performance, our results show that the models generally perform well in detecting and classifying linguistic indicators of category labels used to denote a category, but sometimes struggle to correctly evaluate the associated behaviors and characteristics. Using more few-shot examples within the prompts, significantly improves performance. Model performance increases with size, as Llama-3.3-70B-Instruct and GPT-4 achieve comparable results that surpass those of Mixtral-8x7B-Instruct, GPT-4-mini and Llama-3.1-8B-Instruct.

刻板印象检测大模型偏见语言分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。