首个评估大模型合规判断能力的基准数据集,专测AI法案符合性。
AIReg-Bench: Benchmarking Language Models That Assess AI Regulation Compliance
- 用结构化指令生成120个虚构AI系统技术文档
- 法律专家标注每份文档对欧盟AI法案的违法案
- 为大模型合规评估提供可复现的基准
随着各国推进AI监管,利用大语言模型(LLMs)评估AI系统是否符合人工智能监管法规(AIR)日益受到关注。然而,目前尚无方法对这类评估性能进行基准测试。为此,我们提出AIReg-Bench:首个公开的基准数据集,用于测试LLMs在评估欧盟《人工智能法案》(AIA)合规性方面的能力。该数据集通过两步构建:(1)使用精心设计的提示,由一个LLM生成120个技术文档片段,每个描述一个虚构但合理的AI系统,类似于实际供应商提交的合规证明材料;(2)法律专家对每个样本进行评审并标注其违反《AIA》具体条款的情况。该数据集连同对前沿大模型复现专家标注结果的评估,为理解基于大模型的合规评估工具潜力与局限提供了起点,并建立了后续模型对比的基准。数据集与评估代码已开源:https://github.com/camlsys/aireg-bench。
原文摘要 · Abstract (English)
As governments move to regulate AI, there is growing interest in using Large Language Models (LLMs) to assess whether or not an AI system complies with a given AI Regulation (AIR). However, there is presently no way to benchmark the performance of LLMs at this task. To fill this void, we introduce AIReg-Bench: the first open benchmark dataset designed to test how well LLMs can assess compliance with the EU AI Act (AIA). We created this dataset through a two-step process: (1) by prompting an LLM with carefully structured instructions, we generated 120 technical documentation excerpts (samples), each depicting a fictional, albeit plausible, AI system -- of the kind an AI provider might produce to demonstrate their compliance with AIR; (2) legal experts then reviewed and annotated each sample to indicate whether, and in what way, the AI system described therein violates specific Articles of the AIA. The resulting dataset, together with our evaluation of whether frontier LLMs can reproduce the experts' compliance labels, provides a starting point to understand the opportunities and limitations of LLM-based AIR compliance assessment tools and establishes a benchmark against which subsequent LLMs can be compared. The dataset and evaluation code are available at https://github.com/camlsys/aireg-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。