arXiv:2410.09978cs.CLcs.CY2024-10被引 8

检测大模型新闻摘要的政治偏见,发现对民主党倾向明显。

When Neutral Summaries are not that Neutral: Quantifying Political Neutrality in LLM-Generated News Summaries

  • 用抽象摘要方法量化大模型对五大政治议题的偏见
  • 枪支与医疗议题最大偏见达-9.49%和-6.14%
  • 输出词汇高度趋同,适合关注AI伦理的研究者

在算法塑造社会叙事的时代,探究大模型的政治中立性至关重要。本研究通过抽象文本摘要视角,量化大模型在五项美国重大政治议题——堕胎、枪支管控/权利、医疗、移民和LGBTQ+权利——上的中立性。基于20,344篇新闻文章的语料库,研究发现多个主流大模型存在系统性亲民主党偏见,其中枪支管控与医疗议题的极化差异最大,分别为-9.49%和-6.14%。进一步分析显示,这些敏感议题的模型输出词汇呈现显著收敛:民主党倾向表述间词汇重叠率达55%,共和党倾向达52%。在关键选举临近之际,此发现具有重要意义。

原文摘要 · Abstract (English)

In an era where societal narratives are increasingly shaped by algorithmic curation, investigating the political neutrality of LLMs is an important research question. This study presents a fresh perspective on quantifying the political neutrality of LLMs through the lens of abstractive text summarization of polarizing news articles. We consider five pressing issues in current US politics: abortion, gun control/rights, healthcare, immigration, and LGBTQ+ rights. Via a substantial corpus of 20,344 news articles, our study reveals a consistent trend towards pro-Democratic biases in several well-known LLMs, with gun control and healthcare exhibiting the most pronounced biases (max polarization differences of -9.49% and -6.14%, respectively). Further analysis uncovers a strong convergence in the vocabulary of the LLM outputs for these divisive topics (55% overlap for Democrat-leaning representations, 52% for Republican). Being months away from a US election of consequence, we consider our findings important.

大模型偏见政治中立摘要生成舆情分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。