发现生成式AI内容审核会误删身份相关言论,尤其针对弱势群体。
Identity-related Speech Suppression in Generative AI Content Moderation
- 构建九类身份群体的言论抑制评估基准,测试多个审核系统
- 身份相关文本被错误标记概率高于普通文本,差异显著
- 不同身份受偏见影响不同,如残障内容易被误判为自残
自动化内容审核长期用于识别和过滤用户生成的不当内容,但常错误标记涉及边缘化身份的内容。如今生成式AI也采用此类过滤器,防止生成或展示不当内容。然而,现有研究多关注避免生成不良结果,却忽视了允许适当文本生成的重要性。随着生成式AI在教学、影视等场景中广泛应用,这些技术将允许谁讲故事,又压制谁的声音?本文定义并引入言论抑制的衡量方法,聚焦于多种身份相关文本被内容审核API错误过滤的情况。我们使用传统短文本数据集及两个新引入的生成式AI专用数据集,建立针对九类身份群体的言论抑制评估基准。测试涵盖一个传统与四个生成式AI导向的自动审核服务,结果显示:身份相关言论被错误抑制的概率显著高于其他内容。错误标记原因因身份而异,受刻板印象和文本关联影响——例如,残障相关内容更可能被误标为自残或健康风险,非基督教内容则更易被误标为暴力或仇恨。随着生成式AI广泛应用于创作,亟需关注其对身份相关内容生成的影响。
原文摘要 · Abstract (English)
Automated content moderation has long been used to help identify and filter undesired user-generated content online. But such systems have a history of incorrectly flagging content by and about marginalized identities for removal. Generative AI systems now use such filters to keep undesired generated content from being created by or shown to users. While a lot of focus has been given to making sure such systems do not produce undesired outcomes, considerably less attention has been paid to making sure appropriate text can be generated. From classrooms to Hollywood, as generative AI is increasingly used for creative or expressive text generation, whose stories will these technologies allow to be told, and whose will they suppress? In this paper, we define and introduce measures of speech suppression, focusing on speech related to different identity groups incorrectly filtered by a range of content moderation APIs. Using both short-form, user-generated datasets traditional in content moderation and longer generative AI-focused data, including two datasets we introduce in this work, we create a benchmark for measurement of speech suppression for nine identity groups. Across one traditional and four generative AI-focused automated content moderation services tested, we find that identity-related speech is more likely to be incorrectly suppressed than other speech. We find that reasons for incorrect flagging behavior vary by identity based on stereotypes and text associations, with, e.g., disability-related content more likely to be flagged for self-harm or health-related reasons while non-Christian content is more likely to be flagged as violent or hateful. As generative AI systems are increasingly used for creative work, we urge further attention to how this may impact the creation of identity-related content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。