用三权分立机制让大模型更懂伦理,情绪调控是关键。
A Checks-and-Balances Framework for Context-Aware Ethical AI Alignment
- 三组件分工:执行(生成)、立法(设伦理规则)、司法(判情境)
- 通过情绪映射语言行为,实现精准伦理引导
- 适合关注AI伦理对齐与安全控制的研究者
本文提出一种受三权分立体制启发的伦理对齐框架,用于大型语言模型(LLMs)。框架包含三个独立但交互的组件:作为执行机构的LLM负责知识生成,作为立法机构的DIKE建立伦理约束,作为司法机构的ERIS进行上下文解读。为解决情感调控这一核心挑战,基于心理学理论,设计了自监督学习流程,将情绪映射到语言行为,实现通过情绪调节实现行为精准控制。结合对抗测试,该框架证明了在知识生成、伦理监督和情境理解中保持独立性的同时,可引导语言行为向合乎伦理方向发展。
原文摘要 · Abstract (English)
This paper introduces a checks-and-balances framework for ethical alignment of Large Language Models (LLMs), inspired by three-branch governmental systems. It implements three independent yet interacting components: LLMs as the executive branch for knowledge generation, DIKE as the legislative branch establishing ethical guardrails, and ERIS as the judicial branch for contextual interpretation. Beyond structural separation, we address a fundamental challenge: regulating emotion to shape behaviors. Drawing from psychological theories where managing emotional responses prevents harmful behaviors, we develop a self-supervised learning pipeline that maps emotions to linguistic behaviors, enabling precise behavioral modulation through emotional conditioning. By integrating this approach with adversarial testing, our framework demonstrates how DIKE and ERIS direct linguistic behaviors toward ethical outcomes while preserving independence throughout knowledge generation, ethical oversight, and contextual interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。