arXiv:2508.18245cs.CL2025-08

测试大模型对性别歧视的感知差异,发现其难以反映不同人群的真实看法。

Demographic Biases and Gaps in the Perception of Sexism in Large Language Models

  • 用六类人群标注的推特数据评估模型对性别歧视的识别能力
  • 模型整体表现尚可,但无法准确模拟不同年龄性别群体的感知差异
  • 揭示了模型在多元视角建模上的缺陷,适合关注公平性研究者阅读

大型语言模型(LLMs)在自动检测性别歧视方面展现出潜力,但已有研究显示这些模型存在偏差,尤其对少数群体反映不准确。尽管已采取多种改进措施,该任务仍因主观性强及模型固有偏见而面临挑战。本文利用EXIST 2024推特数据集,该数据集为每条推文提供六种不同人群的标注,评估不同LLMs在性别歧视识别中对各群体感知的模仿程度。同时分析模型中的性别与年龄等人口统计学偏差,并通过统计分析识别出影响判断的关键特征。结果显示,尽管模型在总体层面具备一定识别能力,却未能真实再现不同人口群体间的感知多样性。这表明亟需更精细化校准的模型,以更好反映跨群体视角差异。

原文摘要 · Abstract (English)

The use of Large Language Models (LLMs) has proven to be a tool that could help in the automatic detection of sexism. Previous studies have shown that these models contain biases that do not accurately reflect reality, especially for minority groups. Despite various efforts to improve the detection of sexist content, this task remains a significant challenge due to its subjective nature and the biases present in automated models. We explore the capabilities of different LLMs to detect sexism in social media text using the EXIST 2024 tweet dataset. It includes annotations from six distinct profiles for each tweet, allowing us to evaluate to what extent LLMs can mimic these groups' perceptions in sexism detection. Additionally, we analyze the demographic biases present in the models and conduct a statistical analysis to identify which demographic characteristics (age, gender) contribute most effectively to this task. Our results show that, while LLMs can to some extent detect sexism when considering the overall opinion of populations, they do not accurately replicate the diversity of perceptions among different demographic groups. This highlights the need for better-calibrated models that account for the diversity of perspectives across different populations.

大模型偏见性别歧视社会感知公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。