arXiv:2506.06636cs.CL2025-06ACL被引 14

首个从法律角度评估大模型安全性的基准,揭示顶尖模型仍有明显安全短板。

SafeLawBench: Towards Safe Alignment of Large Language Models

  • 基于法律标准将风险分三级,构建系统化安全评估框架
  • 20个模型平均准确率68.8%,顶级模型最高仅80.5%(多选题)
  • 提出多数投票可提升性能,适合关注模型安全的研究者

随着大语言模型(LLMs)广泛应用,其安全性引发广泛关注。然而,由于现有安全评测标准主观性强,尚无统一评判依据。为此,本文首次从法律视角出发,提出SafeLawBench基准,依据法律标准将安全风险划分为三个等级,建立系统性、全面的评估框架。该基准包含24,860道多选题和1,106个开放域问答任务。我们对2个闭源模型与18个开源模型进行了零样本与少样本提示测试,分析各模型的安全特性、推理稳定性及拒绝行为。结果表明,多数投票机制可提升模型表现。值得注意的是,即使领先SOTA模型如Claude-3.5-Sonnet和GPT-4o,在多选任务上准确率也未超过80.5%,20个模型平均准确率为68.8%。研究呼吁学界更加重视大模型的安全性研究。

原文摘要 · Abstract (English)

With the growing prevalence of large language models (LLMs), the safety of LLMs has raised significant concerns. However, there is still a lack of definitive standards for evaluating their safety due to the subjective nature of current safety benchmarks. To address this gap, we conducted the first exploration of LLMs' safety evaluation from a legal perspective by proposing the SafeLawBench benchmark. SafeLawBench categorizes safety risks into three levels based on legal standards, providing a systematic and comprehensive framework for evaluation. It comprises 24,860 multi-choice questions and 1,106 open-domain question-answering (QA) tasks. Our evaluation included 2 closed-source LLMs and 18 open-source LLMs using zero-shot and few-shot prompting, highlighting the safety features of each model. We also evaluated the LLMs' safety-related reasoning stability and refusal behavior. Additionally, we found that a majority voting mechanism can enhance model performance. Notably, even leading SOTA models like Claude-3.5-Sonnet and GPT-4o have not exceeded 80.5% accuracy in multi-choice tasks on SafeLawBench, while the average accuracy of 20 LLMs remains at 68.8\%. We urge the community to prioritize research on the safety of LLMs.

大模型安全法律评估评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。