RAG会放大文档中的社会偏见,即使生成模型本身较中立。
Evaluating the Effect of Retrieval Augmentation on Social Biases
- 用不同偏见水平的文档测试RAG生成结果
- 三语言四类偏见下,生成文本偏见普遍增强
- 提醒需在部署前评估RAG的社会偏见风险
检索增强生成(RAG)因其能便捷地向大语言模型(LLM)驱动的自然语言生成(NLG)系统注入预训练阶段未见的新知识而广受欢迎。然而,LLM通常携带显著的社会偏见。目前对RAG如何调节这些偏见尚不明确。本文系统研究了RAG系统各组件与生成文本中社会偏见的关系,覆盖英语、日语和中文三种语言,以及性别、种族、年龄和宗教四类社会偏见。基于偏差问答(BBQ)基准数据集,我们评估了在具有不同刻板印象水平的文档集合上,使用多个LLM作为生成器时的偏见表现。结果发现,即使生成模型本身偏见水平较低,文档集合中的偏见仍常在生成响应中被放大。这一发现警示了将RAG用于注入新知识时可能引入的潜在社会偏见,呼吁在实际部署前对RAG应用进行谨慎评估。
原文摘要 · Abstract (English)
Retrieval Augmented Generation (RAG) has gained popularity as a method for conveniently incorporating novel facts that were not seen during the pre-training stage in Large Language Model (LLM)-based Natural Language Generation (NLG) systems. However, LLMs are known to encode significant levels of unfair social biases. The modulation of these biases by RAG in NLG systems is not well understood. In this paper, we systematically study the relationship between the different components of a RAG system and the social biases presented in the text generated across three languages (i.e. English, Japanese and Chinese) and four social bias types (i.e. gender, race, age and religion). Specifically, using the Bias Question Answering (BBQ) benchmark datasets, we evaluate the social biases in RAG responses from document collections with varying levels of stereotypical biases, employing multiple LLMs used as generators. We find that the biases in document collections are often amplified in the generated responses, even when the generating LLM exhibits a low-level of bias. Our findings raise concerns about the use of RAG as a technique for injecting novel facts into NLG systems and call for careful evaluation of potential social biases in RAG applications before their real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。