用法律条文指导AI自我修正,提升合规性与效率。
Statutory AI: Aligning Large Language Models With Legal Norms

- 以法律文本为宪法框架,用思维链分步审查输出
- 在5类违规场景中减少有害内容52%-59%
- 比传统方法快一倍以上,适合政策合规场景
随着人工智能监管框架的发展,确保生成式模型等AI系统符合法律与伦理标准已成为关键任务。现有对齐方法存在局限:如宪法AI依赖人工监督,而广义原则(如人类福祉原则)过于模糊,难以提供可操作的治理指引。为此,我们提出混合方法Statutory AI,利用法律语料库中已有的具体主题条款作为行为规范。该方法通过两阶段思维链提示实现:第一阶段将用户请求分类至预设主题,第二阶段结合对应法律条文分析输出内容。实验使用1000个红队测试提示,在歧视、机密泄露、暴力、欺诈及滥用弱势群体等五类惩罚主题中,Statutory AI使有害内容减少52至59个百分点,较标准宪法AI高出约10个百分点,同时计算时间减少超过50%。
原文摘要 · Abstract (English)
With the increasing development of AI regulatory frameworks, ensuring that artificial intelligence systems, particularly generative models, operate in accordance with legal and ethical standards has become a critical priority. Existing proposals for AI alignment and value-guided behavior, however, face some limitations. Approaches such as Constitutional AI depend on human supervision, while broad normative frameworks like the Good-for-Humanity (GfH) principle may be overly general and ambiguous to provide actionable governance guidance. To overcome these limitations, we propose a hybrid approach called Statutory AI that employs pre-existing human-authored principles drawn from specific themes within a legal corpus. Specifically, Statutory AI uses legal texts as a constitutional framework, enabling AI systems to autonomously critique and revise their outputs according to established norms. It operates in two stages, both using Chain-of-Thought prompting. The first stage classifies the user prompt into one of the identified themes, while the second stage analyzes it in conjunction with relevant articles selected from the legal corpus of that theme. To illustrate the potential of our approach, we conducted an experiment involving 1,000 red-teaming prompts and five penal themes: discrimination, disclosure of confidential information, violence, fraud, and abuse of vulnerable persons. Statutory AI reduced harmful content by 52 to 59 percentage points across tested models, approximately 10 percentage points higher than standard Constitutional AI, while cutting computation time by over 50%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。