RAG通过外部知识降低大模型偏见,但思维链会加重偏见。
Evaluating Social Bias in RAG Systems: When External Context Helps and Reasoning Hurts
- 用外部知识检索缓解大模型刻板印象
- 引入检索内容后13类偏见普遍下降
- 思维链虽提准率却加剧偏见,需新框架
大语言模型中的社会偏见引发严重公平问题。尽管检索增强生成(RAG)架构通过引入外部知识源提升生成能力,仍面临相同偏见挑战。本文在多个检索语料库、大模型及偏见评估数据集上进行大规模实验,涵盖超过13种偏见类型,意外发现RAG能降低偏见。这表明外部上下文有助于抵消刻板印象驱动的预测,通过丰富输出的语境基础提升公平性。为深入理解该现象,我们进一步将思维链(CoT)提示融入RAG,并评估模型推理过程的忠实度。实验显示,随着检索文档提供更多上下文,模型偏见倾向在刻板与反刻板响应间波动。有趣的是,尽管CoT提升了准确性,却导致整体偏见上升,揭示了偏见感知推理框架的必要性。
原文摘要 · Abstract (English)
Social biases inherent in large language models (LLMs) raise significant fairness concerns. Retrieval-Augmented Generation (RAG) architectures, which retrieve external knowledge sources to enhance the generative capabilities of LLMs, remain susceptible to the same bias-related challenges. This work focuses on evaluating and understanding the social bias implications of RAG. Through extensive experiments across various retrieval corpora, LLMs, and bias evaluation datasets, encompassing more than 13 different bias types, we surprisingly observe a reduction in bias in RAG. This suggests that the inclusion of external context can help counteract stereotype-driven predictions, potentially improving fairness by diversifying the contextual grounding of the model's outputs. To better understand this phenomenon, we then explore the model's reasoning process by integrating Chain-of-Thought (CoT) prompting into RAG while assessing the faithfulness of the model's CoT. Our experiments reveal that the model's bias inclinations shift between stereotype and anti-stereotype responses as more contextual information is incorporated from the retrieved documents. Interestingly, we find that while CoT enhances accuracy, contrary to the bias reduction observed with RAG, it increases overall bias across datasets, highlighting the need for bias-aware reasoning frameworks that can mitigate this trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。