多类型偏见叠加会严重误导大模型,现有方法难以应对。
Large Language Models Are Still Misled by Simple Bias Ensembles
- 构建包含多种偏见的综合测试集,模拟真实复杂场景。
- 主流大模型在多偏见环境下性能显著下降,准确率大幅降低。
- 适合关注模型可靠性与公平性的研究人员参考。
随着大语言模型(LLMs)的发展,其对单一简单偏见的鲁棒性有所提升。然而,我们发现多种简单偏见的组合仍会对 LLMs 产生显著负面影响。由于真实世界数据通常受多种偏见共现影响,大模型在临床诊断、法律文件分析等高风险场景中表现不稳定。然而,现有基准测试仅针对单类偏见注入的数据集。为填补这一空白,我们提出一个多重偏见基准,每个样本包含多种偏见类型。实验表明,现有 LLMs 及去偏方法在此基准上表现不佳,凸显了消除复合偏见的挑战。
原文摘要 · Abstract (English)
With the evolution of large language models (LLMs), their robustness against individual simple biases has been enhanced. However, we observe that the ensemble of multiple simple biases still exerts a significant adverse impact on LLMs. Given that real-world data samples are typically confounded by a wide range of biases, LLMs tend to exhibit unstable performance when deployed in high-stakes real-world scenarios such as clinical diagnosis and legal document analysis. However, previous benchmarks are constrained to datasets where each sample is manually injected with only one type of bias. To bridge this gap, we propose a multi-bias benchmark where each sample contains multiple types of biases. Experimental results reveal that existing LLMs and debiasing methods perform poorly on this benchmark, highlighting the challenge of eliminating such compounded biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。