构建极端场景公平性测试基准,发现大模型易受简单恶意指令诱导偏见。
FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
- 设计对抗性提示触发模型偏见,评估其在极端场景下的公平性鲁棒性。
- 对比实验显示传统评测低估了模型内在风险,暴露严重安全隐患。
- 适合关注AI伦理与安全评估的研究者和开发者使用。
大型语言模型(LLMs)的进展显著提升了人机交互体验,但随之而来的社会偏见问题可能引发有害社会影响,亟需严格的安全部署评估。现有基准可能忽略模型内在弱点——即使面对简单的对抗性指令,仍会生成偏见响应。为此,我们提出新基准FLEX(Fairness Benchmark in LLM under Extreme Scenarios),专门测试模型在极端诱导性提示下维持公平性的能力。通过引入放大潜在偏见的提示,全面评估模型鲁棒性。对比实验表明,传统评估可能严重低估模型风险,凸显建立更严格评测标准的必要性,以确保大模型的安全与公平。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have significantly enhanced interactions between users and models. These advancements concurrently underscore the need for rigorous safety evaluations due to the manifestation of social biases, which can lead to harmful societal impacts. Despite these concerns, existing benchmarks may overlook the intrinsic weaknesses of LLMs, which can generate biased responses even with simple adversarial instructions. To address this critical gap, we introduce a new benchmark, Fairness Benchmark in LLM under Extreme Scenarios (FLEX), designed to test whether LLMs can sustain fairness even when exposed to prompts constructed to induce bias. To thoroughly evaluate the robustness of LLMs, we integrate prompts that amplify potential biases into the fairness assessment. Comparative experiments between FLEX and existing benchmarks demonstrate that traditional evaluations may underestimate the inherent risks in models. This highlights the need for more stringent LLM evaluation benchmarks to guarantee safety and fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。