测试大模型对阿拉伯人与西方人的偏见及安全漏洞,发现多数模型存在显著偏见且易被攻击。
Desert Camels and Oil Sheikhs: Arab-Centric Red Teaming of Frontier LLMs
- 构建双数据集评估阿拉伯人偏见与对抗性攻击
- 79%的案例显示对阿拉伯人负面偏见,部分模型攻击成功率超87%
- 尽管优化,GPT-4o仍最易被攻破,Claude 3.5最安全但仍有偏见
大型语言模型(LLMs)广泛应用却存在社会偏见问题。本研究在八个领域(包括女性权利、恐怖主义和反犹主义)考察了模型对阿拉伯人与西方人的偏见,并评估其抵抗偏见传播的能力。为此,我们构建了两个数据集:一个用于评估模型对阿拉伯人与西方人的偏见,另一个用于测试能夸大负面特质的‘越狱’提示(jailbreaks)。评估了六款模型:GPT-4、GPT-4o、LlaMA 3.1(8B & 405B)、Mistral 7B 和 Claude 3.5 Sonnet。结果显示,79%的情况中模型表现出对阿拉伯人的负面偏见,其中 LlaMA 3.1-405B 偏见最严重。越狱测试表明,尽管是优化版本,GPT-4o 最脆弱,其次为 LlaMA 3.1-8B 与 Mistral 7B。除 Claude 外,所有模型在三个类别中的攻击成功率均超过 87%。尽管 Claude 3.5 Sonnet 是最安全的,但在八项类别中有七项仍存在偏见。研究揭示出优化版本也可能存在偏见与安全缺陷,亟需更强的偏见缓解与安全机制。
原文摘要 · Abstract (English)
Large language models (LLMs) are widely used but raise ethical concerns due to embedded social biases. This study examines LLM biases against Arabs versus Westerners across eight domains, including women's rights, terrorism, and anti-Semitism and assesses model resistance to perpetuating these biases. To this end, we create two datasets: one to evaluate LLM bias toward Arabs versus Westerners and another to test model safety against prompts that exaggerate negative traits ("jailbreaks"). We evaluate six LLMs -- GPT-4, GPT-4o, LlaMA 3.1 (8B & 405B), Mistral 7B, and Claude 3.5 Sonnet. We find 79% of cases displaying negative biases toward Arabs, with LlaMA 3.1-405B being the most biased. Our jailbreak tests reveal GPT-4o as the most vulnerable, despite being an optimized version, followed by LlaMA 3.1-8B and Mistral 7B. All LLMs except Claude exhibit attack success rates above 87% in three categories. We also find Claude 3.5 Sonnet the safest, but it still displays biases in seven of eight categories. Despite being an optimized version of GPT4, We find GPT-4o to be more prone to biases and jailbreaks, suggesting optimization flaws. Our findings underscore the pressing need for more robust bias mitigation strategies and strengthened security measures in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。