用少量参数微调,显著降低大模型对酷儿群体的偏见。
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
- 采用低秩适应(LoRA)轻量微调,仅增加不到0.1%参数。
- 在酷儿新闻数据集上微调后,偏见得分最高下降50点,中立率升至36%。
- 适合关注AI公平性、想低成本改进模型偏见的研究者和开发者。
大型语言模型常复现训练语料中的性别与性取向偏见,导致对LGBTQIA+用户的输出边缘化。为缓解此问题,我们评估了两种参数高效微调(PEFT)技术——低秩适应(LoRA)与软提示调优,作为全模型微调的轻量替代方案。基于WinoQueer基准,我们在三款开源LLM中量化偏见,发现基线偏见得分高达98(满分100),其中50代表中立。在精选的QueerNews语料上使用LoRA微调(<0.1%额外参数),偏见得分最多降低50分,中立率从近乎0%提升至最高36%。而软提示调优(10个虚拟标记)效果有限。结果表明,LoRA可在极低计算成本下实现显著公平性提升。我们倡导更广泛采用社区共建的PEFT方法,构建更大规模的酷儿作者语料库,并发展超越WinoQueer的丰富评估体系,配合持续审计以确保模型包容性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) frequently reproduce the gender- and sexual-identity prejudices embedded in their training corpora, leading to outputs that marginalize LGBTQIA+ users. Hence, reducing such biases is of great importance. To achieve this, we evaluate two parameter-efficient fine-tuning (PEFT) techniques - Low-Rank Adaptation (LoRA) and soft-prompt tuning - as lightweight alternatives to full-model fine-tuning for mitigating such biases. Using the WinoQueer benchmark, we quantify bias in three open-source LLMs and observe baseline bias scores reaching up to 98 (out of 100) across a range of queer identities defined by gender and/or sexual orientation, where 50 would indicate neutrality. Fine-tuning with LoRA (< 0.1% additional parameters) on a curated QueerNews corpus reduces those scores by up to 50 points and raises neutrality from virtually 0% to as much as 36%. Soft-prompt tuning (10 virtual tokens) delivers only marginal improvements. These findings show that LoRA can deliver meaningful fairness gains with minimal computation. We advocate broader adoption of community-informed PEFT, the creation of larger queer-authored corpora, and richer evaluation suites beyond WinoQueer, coupled with ongoing audits to keep LLMs inclusive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。