发现大模型安全机制对不同群体拒绝生成内容存在偏见。
Characterizing Selective Refusal Bias in Large Language Models
- 分析模型对不同性别、性取向等群体的拒绝率差异
- 发现针对特定群体的拒绝率显著更高
- 提示安全机制需更公平,适合关注模型伦理的研究者
大型语言模型(LLMs)的安全防护机制旨在防止恶意用户大规模生成有害内容。然而,这些措施可能无意中引入或反映新的偏见,即模型可能对某些人口群体的有害内容请求拒绝,而对其他群体则不拒绝。我们通过分析目标个体和交叉性人口群体的拒绝率、响应类型及拒绝文本长度,研究了这一选择性拒绝偏见。结果表明,在性别、性取向、国籍和宗教属性上均存在选择性拒绝偏见。为此,我们进一步通过间接攻击测试,针对曾被拒绝的群体进行攻击,揭示了潜在的安全隐患。研究强调,需在各类人口群体间实现更公平且稳健的安全防护性能。
原文摘要 · Abstract (English)
Safety guardrails in large language models(LLMs) are developed to prevent malicious users from generating toxic content at a large scale. However, these measures can inadvertently introduce or reflect new biases, as LLMs may refuse to generate harmful content targeting some demographic groups and not others. We explore this selective refusal bias in LLM guardrails through the lens of refusal rates of targeted individual and intersectional demographic groups, types of LLM responses, and length of generated refusals. Our results show evidence of selective refusal bias across gender, sexual orientation, nationality, and religion attributes. This leads us to investigate additional safety implications via an indirect attack, where we target previously refused groups. Our findings emphasize the need for more equitable and robust performance in safety guardrails across demographic groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。