通过微调输入揭示大模型隐性偏见,发现其对少数群体更易出错。
Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation
- 用三步可插拔增强法扰动输入,检测模型偏见
- 在BBQ数据集上,主流大模型偏见率显著上升
- 对研究不足的群体,模型偏差更严重,需扩大研究范围
大型语言模型因训练数据的判别性特征而表现出刻板印象偏见。尽管已有方法试图避免使用刻板信息进行决策,但近期研究表明这些对齐方法脆弱易破。本文提出一种通用的增强框架,包含三个即插即用步骤,适用于多种公平性评估基准。将该方法应用于公平性评估数据集(问答偏见基准BBQ),发现包括最先进的开源与闭源模型在内的大模型,对输入扰动极为敏感,表现出更高的刻板行为倾向。此外,当目标人群为文献中研究较少的群体时,模型更易产生偏见,凸显了扩展公平性与安全性研究以涵盖更多元群体的迫切需求。
原文摘要 · Abstract (English)
Large Language Models have been shown to demonstrate stereotypical biases in their representations and behavior due to the discriminative nature of the data that they have been trained on. Despite significant progress in the development of methods and models that refrain from using stereotypical information in their decision-making, recent work has shown that approaches used for bias alignment are brittle. In this work, we introduce a novel and general augmentation framework that involves three plug-and-play steps and is applicable to a number of fairness evaluation benchmarks. Through application of augmentation to a fairness evaluation dataset (Bias Benchmark for Question Answering (BBQ)), we find that Large Language Models (LLMs), including state-of-the-art open and closed weight models, are susceptible to perturbations to their inputs, showcasing a higher likelihood to behave stereotypically. Furthermore, we find that such models are more likely to have biased behavior in cases where the target demographic belongs to a community less studied by the literature, underlining the need to expand the fairness and safety research to include more diverse communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。