arXiv:2504.07887cs.CLcs.AI2025-04被引 34

用AI当裁判,自动测试大模型抗偏见攻击能力。

Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge

  • 让大模型自评偏见风险,构建可扩展的评测框架。
  • 发现年龄、残疾及交叉偏见最易被诱发,小模型有时比大模型更安全。
  • 适合关注AI公平性与安全性的研究人员和开发者。

大型语言模型(LLMs)在关键社会领域日益普及,但其内部偏见可能加剧刻板印象并破坏公平性,这些偏见源于训练数据的历史不公、语言失衡或对抗性操纵。尽管已有缓解措施,最新研究显示LLMs仍易受诱发偏见的对抗攻击。本文提出一种可扩展的基准评测框架,用于评估模型对对抗性偏见诱导的鲁棒性。方法包括:(i) 在多个任务中系统探测模型在不同社会文化偏见上的表现;(ii) 采用LLM-as-a-Judge方法量化安全得分;(iii) 使用越狱技术揭示安全漏洞。为支持系统化评测,我们发布了名为CLEAR-Bias的偏见相关提示数据集。分析发现DeepSeek V3是最佳裁判模型,偏见韧性分布不均,其中年龄、残疾及交叉偏见最为突出。部分小型模型的安全性优于大型模型,表明训练与架构可能比规模更重要。然而,无一模型能完全抵御对抗性诱导,使用低资源语言或拒绝抑制的越狱攻击在各类模型中均有效。此外,连续世代模型呈现轻微安全性提升,而医疗领域微调模型普遍不如通用模型安全。

原文摘要 · Abstract (English)

The growing integration of Large Language Models (LLMs) into critical societal domains has raised concerns about embedded biases that can perpetuate stereotypes and undermine fairness. Such biases may stem from historical inequalities in training data, linguistic imbalances, or adversarial manipulation. Despite mitigation efforts, recent studies show that LLMs remain vulnerable to adversarial attacks that elicit biased outputs. This work proposes a scalable benchmarking framework to assess LLM robustness to adversarial bias elicitation. Our methodology involves: (i) systematically probing models across multiple tasks targeting diverse sociocultural biases, (ii) quantifying robustness through safety scores using an LLM-as-a-Judge approach, and (iii) employing jailbreak techniques to reveal safety vulnerabilities. To facilitate systematic benchmarking, we release a curated dataset of bias-related prompts, named CLEAR-Bias. Our analysis, identifying DeepSeek V3 as the most reliable judge LLM, reveals that bias resilience is uneven, with age, disability, and intersectional biases among the most prominent. Some small models outperform larger ones in safety, suggesting that training and architecture may matter more than scale. However, no model is fully robust to adversarial elicitation, with jailbreak attacks using low-resource languages or refusal suppression proving effective across model families. We also find that successive LLM generations exhibit slight safety gains, while models fine-tuned for the medical domain tend to be less safe than their general-purpose counterparts.

偏见检测大模型安全自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。