语言模型过滤机制过度删除边缘群体相关表述,造成知识性抹除。
Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails

- 基于关键词黑名单的自动过滤与防御机制主导决策
- 跨系统对边缘群体(如跨性别者、女性、中美洲人)过度标记
- 人工标注显示90%以上被删内容实为重要社会表达,当前系统识别不足
现代语言模型依赖预训练过滤器清理训练数据中的不当内容,并在部署时使用推理阶段的防护机制抑制不良输出。本文审计了四个预训练过滤器和三个推理阶段防护机制,在包含性别与地区起源提及的Common Crawl语句上进行分析,同时结合500条人工标注样本。结果表明,过滤与防护决策高度依赖黑名单式词汇线索,却常遗漏隐私信息或明确仇恨言论。同时,边缘群体——尤其是跨性别者、女性及中美洲人——被显著过度标记。相比之下,人工标注者认为88.5%被过滤标记的内容和91.3%被防护机制标记的内容应保留,常能识别出当前系统未能捕捉到的内容移除带来的表征性伤害。整体揭示了一种认知性抹除现象:边缘群体相关表述在预训练前即遭不成比例删除,并在推理阶段再次被压制。
原文摘要 · Abstract (English)
Modern language models rely on pretraining filters to remove undesirable content from training corpora and inference-time guardrails to suppress undesirable outputs during deployment. In this paper, we examine how these filtering and moderation decisions produce forms of epistemic erasure and reveal tensions both across automated systems and between these systems and human judgment. We audit four pretraining filters and three inference-time guardrails on Common Crawl sentences containing gender and regional-origin mentions, together with a manually annotated subset of 500 sentences. Our analysis shows that filtering and guardrail decisions are strongly associated with blocklist-based lexical cues, while frequently failing to flag content containing private information or explicit hate speech. At the same time, marginalized groups, particularly transgender people, women, and Central Americans, are significantly over-flagged across systems. Human annotators, by contrast, would retain 88.5\% of filter-flagged and 91.3\% of guardrail-flagged content, often recognizing representational harms arising from tensions of content removal that current systems fail to capture. Taken together, our findings document a form of epistemic erasure in which mentions of marginalized groups are disproportionately removed before pretraining and additionally suppressed again at inference time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。