用宪法式AI塑造更可控的AI助手人格,提升对话质量与一致性。
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
- 基于宪法AI和合成反思数据,系统化训练助手人格
- 多个人格(如幽默、关怀)下生成更连贯真实,抗对抗性提示更强
- 保持模型通用能力,适合需定制化交互体验的场景
现代聊天机器人大型语言模型生成的AI助手人格,影响其行为表现、价值观及伦理倾向,进而决定交互质量、感知智能水平以及与开发者和用户意图的一致性。这种人格塑造过程称为角色训练,是工业界后训练的关键环节,但在学术研究中仍几乎未被探讨。本文首次公开实现角色训练方法,采用宪法AI并结合新型合成内省数据流水线,相比约束系统提示或激活引导等方法,能更有效、更可控地塑造助手人格。我们对三款主流开源权重模型进行微调,使用11种示例人格(如幽默、深情关怀甚至恶意)。为评估效果,提出一种基于揭示偏好的分析方法,发现人格变化显著且整体一致。结果表明,该方法在对抗性提示下更具鲁棒性,同时生成内容更连贯自然。最后,微调对通用能力影响极小,如在标准基准测试中表现基本不变。完整后训练流程已开源,代码见 https://github.com/maiush/OpenCharacterTraining。
原文摘要 · Abstract (English)
The character of the "AI assistant" persona generated by modern chatbot large language models influences both surface-level behavior and apparent values, beliefs, and ethics. These all affect interaction quality, perceived intelligence, and alignment with both developer and user intentions. The shaping of this persona, known as character training, is a critical component of industry post-training, yet remains effectively unstudied in the academic literature. We introduce the first open implementation of character training, leveraging Constitutional AI and a new data pipeline using synthetic introspective data to shape the assistant persona in a more effective and controlled manner than alternatives such as constraining system prompts or activation steering. Specifically, we fine-tune three popular open-weights models using 11 example personas, such as humorous, deeply caring, or even malevolent. To track the effects of our approach, we introduce a method which analyzes revealed preferences, uncovering clear and holistic changes in character. We find these changes are more robust to adversarial prompting than the above two alternatives, while also leading to more coherent and realistic generations. Finally, we demonstrate this fine-tuning has little to no effect on general capabilities as measured by common benchmarks. We describe and open-source our full post-training method, the implementation of which can be found at https://github.com/maiush/OpenCharacterTraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。