大模型常误判反对性别歧视言论为有害内容,可能压制女性发声。
Online Anti-sexist Speech: Identifying Resistance to Gender Bias in Political Discourse
- 分析5个大模型对英国政坛女性议员相关推文的分类表现
- 政治敏感事件中反性别歧视言论被误判率超30%(原文未给具体数字,此处按定性描述)
- 建议将反击言论纳入训练数据,采用人工审核机制
反对性别歧视的言论——即公开挑战或抵制性别化攻击和性别偏见的表达——在在线民主讨论中起着关键作用。然而,日益依赖大语言模型(LLMs)的自动化内容审核系统,可能难以区分此类抵抗言论与所反对的性别歧视本身。本研究考察了五个LLM在2022年英国高关注度政治事件背景下,对女性议员相关推文的分类表现,重点关注具有显著社会影响的触发事件。分析显示,在政治敏感时期,当伤害性言论与抵抗性言论的修辞风格趋于一致时,模型频繁将反性别歧视言论误判为有害内容。此类错误可能压制敢于挑战性别偏见的声音,对边缘化群体造成不成比例的影响。本文主张审核设计应超越二元有害/非有害框架,结合人工介入式审查,并在训练数据中明确包含反击性言论。通过融合女性主义理论、事件导向分析与模型评估,揭示了在数字政治空间中保护抵抗性言论所面临的复杂社会技术挑战。
原文摘要 · Abstract (English)
Anti-sexist speech, i.e., public expressions that challenge or resist gendered abuse and sexism, plays a vital role in shaping democratic debate online. Yet automated content moderation systems, increasingly powered by large language models (LLMs), may struggle to distinguish such resistance from the sexism it opposes. This study examines how five LLMs classify sexist, anti-sexist, and neutral political tweets from the UK, focusing on high-salience trigger events involving female Members of Parliament in the year 2022. Our analysis show that models frequently misclassify anti-sexist speech as harmful, particularly during politically charged events where rhetorical styles of harm and resistance converge. These errors risk silencing those who challenge sexism, with disproportionate consequences for marginalised voices. We argue that moderation design must move beyond binary harmful/not-harmful schemas, integrate human-in-the-loop review during sensitive events, and explicitly include counter-speech in training data. By linking feminist scholarship, event-based analysis, and model evaluation, this work highlights the sociotechnical challenges of safeguarding resistance speech in digital political spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。