arXiv:2603.13933cs.CL2026-03ACL被引 1

构建10万+真实合规案例数据集,助力大模型安全评测

OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset

  • 基于网页搜索代理采集多领域权威规则与真实案例
  • 涵盖74项法规、1.3万条规则、10.6万例真实合规场景
  • 为大模型安全评测提供可量化的现实基准,适合安全研究者

确保大语言模型(LLMs)的安全与合规至关重要。然而,现有安全数据集常依赖随意分类生成,缺乏真实世界中基于规则的案例,难以有效保障模型安全。本文从合规视角构建全面的安全数据集。通过强大网页搜索代理,我们收集了来自多领域权威来源的规则驱动型真实案例数据集OmniCompliance-100K。数据集覆盖74项法规与政策,涵盖安全隐私、内容安全、用户数据隐私、金融安全、医疗设备风险管理、教育诚信及基本人权保护等多个领域。全集包含12,985条独立规则和106,009个相关真实合规案例。分析显示规则与案例间存在强一致性。我们进一步开展广泛基准测试,评估不同规模先进大模型的安全与合规能力。实验揭示多个重要发现,为未来大模型安全研究提供宝贵洞见。

原文摘要 · Abstract (English)

Ensuring the safety and compliance of large language models (LLMs) is of paramount importance. However, existing LLM safety datasets often rely on ad-hoc taxonomies for data generation and suffer from a significant shortage of rule-grounded, real-world cases that are essential for robustly protecting LLMs. In this work, we address this critical gap by constructing a comprehensive safety dataset from a compliance perspective. Using a powerful web-searching agent, we collect a rule-grounded, real-world case dataset OmniCompliance-100K, sourced from multi-domain authoritative references. The dataset spans 74 regulations and policies across a wide range of domains, including security and privacy regulations, content safety and user data privacy policies from leading AI companies and social media platforms, financial security requirements, medical device risk management standards, educational integrity guidelines, and protections of fundamental human rights. In total, our dataset contains 12,985 distinct rules and 106,009 associated real-world compliance cases. Our analysis confirms a strong alignment between the rules and their corresponding cases. We further conduct extensive benchmarking experiments to evaluate the safety and compliance capabilities of advanced LLMs across different model scales. Our experiments reveal several interesting findings that have great potential to offer valuable insights for future LLM safety research.

大模型安全合规数据集真实案例规则对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。