arXiv:2603.20572cs.LG2026-03中稿 · Transactions on Ma…

构建首个基于法律体系的犯罪行为评估基准,测试大模型对76类犯罪的防御能力

LJ-Bench: Ontology-Based Benchmark for U.S. Crime

  • 基于美国刑法典与加州法构建犯罪概念本体,实现分类标准化
  • 覆盖76种犯罪类型,发现大模型对危害社会的行为更易受攻击
  • 开源数据集与代码,助力安全可控的大模型研发

大型语言模型(LLMs)在面对大量非法查询时可能产生有害信息,已成为重大隐患。现有评测基准仅覆盖少数非法行为且缺乏法律依据。本文基于《模范刑法典》(Model Penal Code)构建犯罪概念本体,并以加州法律为实例,形成结构化知识体系,支撑首个全面的犯罪行为评测基准 LJ-Bench。该基准涵盖76种按分类组织的犯罪类型,支持系统性评估多种攻击形式。结果揭示:大模型对造成社会危害的攻击比针对个人的攻击更具脆弱性。本研究旨在推动更鲁棒、可信的 LLM 发展。LJ-Bench 基准、LJ-Ontology 及实验实现均已公开于 https://github.com/AndreaTseng/LJ-Bench。

原文摘要 · Abstract (English)

The potential of Large Language Models (LLMs) to provide harmful information remains a significant concern due to the vast breadth of illegal queries they may encounter. Unfortunately, existing benchmarks only focus on a handful types of illegal activities, and are not grounded in legal works. In this work, we introduce an ontology of crime-related concepts grounded in the legal frameworks of Model Panel Code, which serves as an influential reference for criminal law and has been adopted by many U.S. states, and instantiated using Californian Law. This structured knowledge forms the foundation for LJ-Bench, the first comprehensive benchmark designed to evaluate LLM robustness against a wide range of illegal activities. Spanning 76 distinct crime types organized taxonomically, LJ-Bench enables systematic assessment of diverse attacks, revealing valuable insights into LLM vulnerabilities across various crime categories: LLMs exhibit heightened susceptibility to attacks targeting societal harm rather than those directly impacting individuals. Our benchmark aims to facilitate the development of more robust and trustworthy LLMs. The LJ-Bench benchmark and LJ-Ontology, along with experiments implementation for reproducibility are publicly available at https://github.com/AndreaTseng/LJ-Bench.

大模型安全犯罪评测本体构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。