用法律框架提升大模型安全,让AI行为合规更可靠
Safety Compliance: Rethinking LLM Safety Reasoning through the Lens of Compliance
- 以欧盟人工智能法案和GDPR为标准构建安全评测基准
- 通过强化学习使大模型在法律合规性上提升10%以上
- 适合关注AI监管与安全对齐的研究者和开发者
大型语言模型(LLM)的广泛应用凸显了其安全性的关键重要性。然而,现有安全方法依赖非系统化的分类体系,难以应对现代LLM复杂多变的行为。本文提出从法律合规视角重新定义大模型安全,即‘安全合规’。我们以欧盟人工智能法案(EU AI Act)和通用数据保护条例(GDPR)为核心法律框架,构建了一个基于真实法律条文生成的安全场景新基准。进一步,采用分组策略优化(GRPO)对Qwen3-8B进行对齐训练,构建出名为Compliance Reasoner的安全推理器,使其有效遵循法律标准以降低安全风险。全面实验表明,该推理器在新基准上表现优异,对欧盟人工智能法案平均提升10.45%,对GDPR平均提升11.85%。
原文摘要 · Abstract (English)
The proliferation of Large Language Models (LLMs) has demonstrated remarkable capabilities, elevating the critical importance of LLM safety. However, existing safety methods rely on ad-hoc taxonomy and lack a rigorous, systematic protection, failing to ensure safety for the nuanced and complex behaviors of modern LLM systems. To address this problem, we solve LLM safety from legal compliance perspectives, named safety compliance. In this work, we posit relevant established legal frameworks as safety standards for defining and measuring safety compliance, including the EU AI Act and GDPR, which serve as core legal frameworks for AI safety and data security in Europe. To bridge the gap between LLM safety and legal compliance, we first develop a new benchmark for safety compliance by generating realistic LLM safety scenarios seeded with legal statutes. Subsequently, we align Qwen3-8B using Group Policy Optimization (GRPO) to construct a safety reasoner, Compliance Reasoner, which effectively aligns LLMs with legal standards to mitigate safety risks. Our comprehensive experiments demonstrate that the Compliance Reasoner achieves superior performance on the new benchmark, with average improvements of +10.45% for the EU AI Act and +11.85% for GDPR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。