测试数据过滤策略发现:减少有害内容反而加剧弱势群体缺失。
What Are They Filtering Out? An Experimental Benchmark of Filtering Strategies for Harm Reduction in Pretraining Datasets
- 构建实验基准,系统评估55篇论文中的过滤策略
- 过滤虽降有害内容,但使歧视性群体在数据中更少出现
- 提醒从业者警惕过滤带来的隐性偏见放大
数据过滤是构建安全大语言模型的关键环节,可有效移除预训练数据中的有害内容。然而,现有研究缺乏对这些策略对歧视性弱势群体实际影响的系统评估。本文针对55份英文大模型技术报告,梳理现有过滤策略,并设计实验验证其对弱势群体的影响。结果表明,尽管过滤策略能有效降低文档中的有害内容,但同时导致弱势群体在数据集中被进一步边缘化,出现显著的代表性不足。本研究揭示了当前过滤方法的潜在风险,强调需在安全性与公平性之间权衡。
原文摘要 · Abstract (English)
Data filtering strategies are a crucial component to develop safe Large Language Models (LLM), since they support the removal of harmful contents from pretraining datasets. There is a lack of research on the actual impact of these strategies on vulnerable groups to discrimination, though, and their effectiveness has not been yet systematically addressed. In this paper we present a benchmark study of data filtering strategies for harm reduction aimed at providing a systematic evaluation on these approaches. We provide an overview $55$ technical reports of English LMs and LLMs to identify the existing filtering strategies in literature and implement an experimental setting to test their impact against vulnerable groups. Our results show that the positive impact that strategies have in reducing harmful contents from documents has the side effect of increasing the underrepresentation of vulnerable groups to discrimination in datasets. WARNING: the paper could contain racist, sexist, violent, and generally offensive contents
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。