arXiv:2608.03627cs.AI2026-08中稿 · KDD

LLM检测假新闻时会因发言者性别标签不同而产生偏差,影响判断可靠性。

Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection

  • 通过替换发言者职业头衔性别特征,测试模型对同一内容的判断差异。
  • 超35%的判断结果因性别呈现变化,男性发言者更易被质疑。
  • 首次揭示模型存在系统性性别偏见,适合关注AI公平性的研究者阅读。

大型语言模型(LLMs)在自动事实核查中的应用日益广泛,但其在该场景下的性别偏见问题仍缺乏系统研究。本研究首次基于真实数据,系统考察了LLM在假新闻检测中的性别偏见。我们对LIAR基准数据集进行扩展,为每条陈述添加中性、男性和女性三种职业头衔变体,以检验真伪判断是否仅因性别呈现而改变。评估了六种前沿LLM,涵盖多个偏见与公平性指标。所有模型均表现出性别敏感性:9.79%-35.13%的陈述在三类变体中出现不一致标签,男性-女性对比下翻转率高达6.5%-23.6%。识别出两种主要偏见表现:不稳定性(判断不一致)和方向性(系统性偏好)。其中五种模型显示出显著的方向性效应,最强表现为对男性发言者的系统性怀疑。研究结果表明,性别偏见损害了基于LLM的假新闻检测的可靠性和公平性,凸显了开展偏见感知评估与缓解策略的重要性。所构建的增强数据集已公开发布,以支持后续研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used for automated fact-checking, yet their susceptibility to gender bias in this context remains underexplored. This study presents the first systematic investigation of gender bias in LLM-based fake news detection using real-world data. We augment the LIAR benchmark with three gender variants of speaker job titles (Neutral, Male, Female) for each statement to test whether veracity judgments vary solely based on gender presentation. Six state-of-the-art LLMs are evaluated across multiple bias and fairness metrics. All models exhibit gender sensitivity: 9.79%-35.13% of statements receive inconsistent labels across the three variants, with Male-Female comparisons showing 6.5%-23.6% flip rates. Two primary bias manifestations are identified: instability (inconsistent judgments) and directionality (systematic favoritism). Five models show statistically significant directional effects, with the strongest effects displaying male-skeptic patterns. These findings demonstrate that gender bias undermines both reliability and fairness in LLM-based fake news detection, highlighting the need for bias-aware evaluation and mitigation strategies. The augmented dataset is publicly released to support future research.

假新闻检测性别偏见大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。