RAG虽能防幻觉,却可能悄悄引入偏见,连警惕用户也难幸免。
No Free Lunch: Retrieval-Augmented Generation Undermines Fairness in LLMs, Even for Vigilant Users
- 从用户对公平性的认知出发,构建三级偏见影响模型。
- 即使使用完全去偏的数据集,RAG仍会产生偏见输出。
- 揭示现有对齐方法在RAG中失效,呼吁新防护机制。
检索增强生成(RAG)被广泛采用,因其在减少大语言模型(LLM)幻觉和提升领域特定生成能力方面兼具有效性与成本效益。然而,这种有效性与成本效益是否真是‘免费午餐’?本研究从用户公平性意识角度出发,提出一个实用的三级威胁模型,考察不同用户公平性意识水平下对外部数据集的公平性审查程度。我们使用未去偏、部分去偏和完全去偏的数据集,系统评估了RAG的公平性影响。实验表明,无需微调或重新训练,仅通过RAG即可轻易破坏公平性对齐。即便使用完全去偏且看似无偏的外部数据集,RAG仍可能导致偏见输出。研究结果凸显了当前对齐方法在基于RAG的LLM中的局限性,并强调亟需新的策略以保障公平性。本文提出潜在缓解方案,并呼吁进一步研究以开发鲁棒的公平性防护机制。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) is widely adopted for its effectiveness and cost-efficiency in mitigating hallucinations and enhancing the domain-specific generation capabilities of large language models (LLMs). However, is this effectiveness and cost-efficiency truly a free lunch? In this study, we comprehensively investigate the fairness costs associated with RAG by proposing a practical three-level threat model from the perspective of user awareness of fairness. Specifically, varying levels of user fairness awareness result in different degrees of fairness censorship on the external dataset. We examine the fairness implications of RAG using uncensored, partially censored, and fully censored datasets. Our experiments demonstrate that fairness alignment can be easily undermined through RAG without the need for fine-tuning or retraining. Even with fully censored and supposedly unbiased external datasets, RAG can lead to biased outputs. Our findings underscore the limitations of current alignment methods in the context of RAG-based LLMs and highlight the urgent need for new strategies to ensure fairness. We propose potential mitigations and call for further research to develop robust fairness safeguards in RAG-based LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。