arXiv:2501.14940cs.CLcs.AI2025-01ICML被引 23

提出首个考虑上下文的安全性评测基准,让大模型更懂何时该拒绝、何时不该拒绝。

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

  • 基于上下文完整性理论为问题设计真实场景背景
  • 发现上下文显著影响人类判断(p<0.0001)
  • 揭示商业模型在安全语境下仍频繁误拒,适合安全评估研究者

将大语言模型(LLMs)与人类价值观对齐对于其安全部署和广泛应用至关重要。当前的LLM安全性评测多仅关注对单个问题的拒绝行为,忽视了问题发生的具体上下文,可能导致在安全情境下错误拒绝用户请求,降低用户体验。为此,我们提出CASE-Bench——一个融入上下文的可感知安全性评测基准。该基准基于上下文完整性理论,为分类后的查询分配明确且形式化的上下文。此外,不同于以往研究仅依赖少数标注者的多数投票,我们通过功效分析确定了足够数量的标注者,以确保实验条件间差异的统计显著性。我们在多种开源及商用大模型上使用CASE-Bench进行广泛分析,发现上下文对人类判断具有显著且强烈的影响(z检验,p<0.0001),凸显了在安全性评估中引入上下文的必要性。我们还发现,人类判断与模型响应存在明显不一致,尤其在安全上下文中,商用模型表现不佳。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety benchmarks often focus solely on the refusal of individual problematic queries, which overlooks the importance of the context where the query occurs and may cause undesired refusal of queries under safe contexts that diminish user experience. Addressing this gap, we introduce CASE-Bench, a Context-Aware SafEty Benchmark that integrates context into safety assessments of LLMs. CASE-Bench assigns distinct, formally described contexts to categorized queries based on Contextual Integrity theory. Additionally, in contrast to previous studies which mainly rely on majority voting from just a few annotators, we recruited a sufficient number of annotators necessary to ensure the detection of statistically significant differences among the experimental conditions based on power analysis. Our extensive analysis using CASE-Bench on various open-source and commercial LLMs reveals a substantial and significant influence of context on human judgments (p<0.0001 from a z-test), underscoring the necessity of context in safety evaluations. We also identify notable mismatches between human judgments and LLM responses, particularly in commercial models within safe contexts.

安全性评测上下文感知大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。