让大模型自动遵守复杂规则,提升伦理与合规性。
ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code
- 结合规则评分与GRPO优化,实现多维度政策对齐。
- 医疗助手案例中,领域遵循度提升70.9%。
- 适合需高伦理合规性的全球应用部署。
本文提出ArGen(生成式AI的自主调节框架),用于将大语言模型(LLMs)与涵盖伦理原则、操作安全协议及监管合规标准的可配置、机器可读规则对齐。不同于仅基于偏好的对齐方式,ArGen通过原则驱动的自动化奖励评分、组相对策略优化(GRPO)以及受开放策略代理(OPA)启发的治理层,实现多维度政策合规。为验证其在复杂文化价值体系下的可行性,论文以印度教伦理(如不害、正法)为基础,构建医疗AI助手案例,结果显示其在领域范围遵循度上相较基线提升70.9%。开源实现表明,ArGen为可监管、技术可靠、伦理稳健且可验证合规的AI系统提供了可行路径,适用于多样化的全球应用场景。
原文摘要 · Abstract (English)
This paper introduces ArGen (Auto-Regulation of Generative AI systems), a framework for aligning Large Language Models (LLMs) with complex sets of configurable, machine-readable rules spanning ethical principles, operational safety protocols, and regulatory compliance standards. Moving beyond just preference-based alignment, ArGen is designed to ensure LLMs adhere to these multifaceted policies through a novel synthesis of principle-based automated reward scoring, Group Relative Policy Optimisation (GRPO), and an Open Policy Agent (OPA) inspired governance layer. This approach provides the technical foundation for achieving and demonstrating compliance with diverse and nuanced governance requirements. To showcase the framework's capability to operationalize a deeply nuanced and culturally-specific value system, we present an in-depth case study: the development of a medical AI assistant guided by principles from Dharmic ethics (such as Ahimsa and Dharma), as derived from texts like the Bhagavad Gita. This challenging application demonstrates ArGen's adaptability, achieving a 70.9% improvement in domain-scope adherence over the baseline. Through our open-source repository, we show that ArGen's methodology offers a path to 'Governable Al' systems that are technically proficient, ethically robust, and verifiably compliant for safe deployment in diverse global contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。