为智能体设计个性化监督机制,让AI行为更符合不同人的价值观和安全要求。
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values
- 引入‘超我’代理,通过用户自选的规则集动态引导AI决策。
- 在HarmBench等测试中,有害输出减少高达98.3%,拒绝率接近100%。
- 适用于需要高度个性化与安全性的AI系统,如医疗、金融等领域。
具备自主规划与行动能力的智能体系统在多个领域展现出巨大潜力,但其实际部署受限于行为对多样人类价值观、复杂安全需求及合规要求的对齐难题。现有对齐方法在提供个性化上下文时易引发虚构或效率低下问题。本文提出一种新型解决方案:一个名为‘超我’的个性化监督代理,通过引用用户选定的‘信条宪章’(Creed Constitutions)动态引导智能体规划,支持可调适的遵守程度以匹配不可妥协的价值观。实时合规验证器在执行前将计划与这些宪章及通用伦理底线进行比对。我们构建了完整系统,包括原型宪章共享门户,并通过模型上下文协议(MCP)成功集成第三方模型。综合基准测试(HarmBench、AgentHarm)显示,该超我代理显著降低有害输出——在主流大模型(如Gemini 2.5 Flash、GPT-4o)上实现最高达98.3%的有害得分降幅,且在AgentHarm有害数据集上对Claude Sonnet 4的拒绝率达100%。该方法大幅简化个性化对齐过程,使智能体更可靠地契合个体与文化背景,同时提升安全性。研究概述与示例见 https://superego.creed.space。
原文摘要 · Abstract (English)
Agentic AI systems, possessing capabilities for autonomous planning and action, show great potential across diverse domains. However, their practical deployment is hindered by challenges in aligning their behavior with varied human values, complex safety requirements, and specific compliance needs. Existing alignment methodologies often falter when faced with the complex task of providing personalized context without inducing confabulation or operational inefficiencies. This paper introduces a novel solution: a 'superego' agent, designed as a personalized oversight mechanism for agentic AI. This system dynamically steers AI planning by referencing user-selected 'Creed Constitutions' encapsulating diverse rule sets -- with adjustable adherence levels to fit non-negotiable values. A real-time compliance enforcer validates plans against these constitutions and a universal ethical floor before execution. We present a functional system, including a demonstration interface with a prototypical constitution-sharing portal, and successful integration with third-party models via the Model Context Protocol (MCP). Comprehensive benchmark evaluations (HarmBench, AgentHarm) demonstrate that our Superego agent dramatically reduces harmful outputs -- achieving up to a 98.3% harm score reduction and near-perfect refusal rates (e.g., 100% with Claude Sonnet 4 on AgentHarm's harmful set) for leading LLMs like Gemini 2.5 Flash and GPT-4o. This approach substantially simplifies personalized AI alignment, rendering agentic systems more reliably attuned to individual and cultural contexts, while also enabling substantial safety improvements. An overview on this research with examples is available at https://superego.creed.space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。