通过构造对话暴露大模型隐藏的偏见,发现其难以拒绝后续歧视性提问。
CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
- 设计对抗性对话场景,诱导模型暴露社会偏见。
- 11个模型在6类群体上均出现偏见放大,多数无法拒绝歧视性追问。
- 适合关注AI伦理、安全评估的研究者和从业者。
随着模型构建技术进步,大型语言模型(LLMs)在标准安全测试中表现良好,但在实际对话中仍可能流露有害行为,如种族主义言论。为系统分析此现象,我们提出CoBia——一种轻量级对抗攻击工具,通过构造特定对话情境,诱导模型发表针对性别、种族、宗教、国籍、性取向等六类社会人口特征的偏见言论,并评估其能否自我纠正及拒绝后续偏见问题。我们在11个开源与专有模型上进行测试,采用主流的基于LLM的偏见度量指标,并与人类判断对比以评估模型可靠性与对齐程度。结果表明,有意构造的对话能可靠揭示偏见放大现象,且多数模型无法拒绝后续歧视性提问。该压力测试方法可有效挖掘深层嵌入的交互式偏见。代码与数据集已公开于https://github.com/nafisenik/CoBia。
原文摘要 · Abstract (English)
Improvements in model construction, including fortified safety guardrails, allow Large language models (LLMs) to increasingly pass standard safety checks. However, LLMs sometimes slip into revealing harmful behavior, such as expressing racist viewpoints, during conversations. To analyze this systematically, we introduce CoBia, a suite of lightweight adversarial attacks that allow us to refine the scope of conditions under which LLMs depart from normative or ethical behavior in conversations. CoBia creates a constructed conversation where the model utters a biased claim about a social group. We then evaluate whether the model can recover from the fabricated bias claim and reject biased follow-up questions. We evaluate 11 open-source as well as proprietary LLMs for their outputs related to six socio-demographic categories that are relevant to individual safety and fair treatment, i.e., gender, race, religion, nationality, sex orientation, and others. Our evaluation is based on established LLM-based bias metrics, and we compare the results against human judgments to scope out the LLMs' reliability and alignment. The results suggest that purposefully constructed conversations reliably reveal bias amplification and that LLMs often fail to reject biased follow-up questions during dialogue. This form of stress-testing highlights deeply embedded biases that can be surfaced through interaction. Code and artifacts are available at https://github.com/nafisenik/CoBia.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。