arXiv:2603.09435cs.AI2026-03

构建可复现的NLP与RAG系统评测数据集,助力欧盟AI法案合规性评估

AI Act Evaluation Benchmark: An Open, Transparent, and Reproducible Evaluation Dataset for NLP and RAG Systems

  • 基于大模型与领域知识生成高相关性法规场景任务
  • 覆盖风险分级、条款检索等4类任务,支持自动化评估
  • 适合监管科技、AI合规研究者使用

AI在公共与社会领域的快速部署,推动了对监管标准合规性的迫切需求。欧盟《AI法案》成为关键监管框架,但现有评估工具受限于资源匮乏,常依赖易出错的人工判断,且难以覆盖法规中未明确定义的风险边界(如‘有限’‘最低’情形)。本文提出一种开放、透明、可复现的方法,构建面向NLP与RAG系统的评测数据集,涵盖欧盟AI法案的风险等级分类、条款检索、义务生成与问答任务。数据以机器可读格式组织,利用大语言模型结合领域知识生成真实场景与对应任务。该方法实现了高文档相关性的可信生成,有效解决模糊风险边界问题。实验表明,基于RAG的解决方案在禁止类和高风险场景中分别取得0.87和0.85的F1分数,验证了数据集的有效性。

原文摘要 · Abstract (English)

The rapid rollout of AI in heterogeneous public and societal sectors has subsequently escalated the need for compliance with regulatory standards and frameworks. The EU AI Act has emerged as a landmark in the regulatory landscape. The development of solutions that elicit the level of AI systems' compliance with such standards is often limited by the lack of resources, hindering the semi-automated or automated evaluation of their performance. This generates the need for manual work, which is often error-prone, resource-limited or limited to cases not clearly described by the regulation. This paper presents an open, transparent, and reproducible method of creating a resource that facilitates the evaluation of NLP models with a strong focus on RAG systems. We have developed a dataset that contain the tasks of risk-level classification, article retrieval, obligation generation, and question-answering for the EU AI Act. The dataset files are in a machine-to-machine appropriate format. To generate the files, we utilise domain knowledge as an exegetical basis, combining with the processing and reasoning power of large language models to generate scenarios along with the respective tasks. Our methodology demonstrates a way to harness language models for grounded generation with high document relevancy. Besides, we overcome limitations such as navigating the decision boundaries of risk-levels that are not explicitly defined within the EU AI Act, such as limited and minimal cases. Finally, we demonstrate our dataset's effectiveness by evaluating a RAG-based solution that reaches 0.87 and 0.85 F1-score for prohibited and high-risk scenarios.

AI合规RAG评测基准自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。