提出一套可验证的宪法型AI对齐框架,提升模型安全性与人类偏好一致性。
C3AI: Crafting and Evaluating Constitutions for Constitutional AI
- 从心理学和人工智能中筛选正向行为原则,构建更有效的对齐宪法。
- 改进后的宪法在保持推理能力的同时,显著提升安全指标(未给出具体数值)。
- 发现模型对负面原则执行良好,但正向原则难以遵循,揭示设计与实际的差距。
宪法型AI(CAI)通过宪法指导大模型行为,但如何选择最有效的原则仍是个开放问题。本文提出C3AI框架(构建宪法型AI模型),兼具两大功能:(1) 在微调前筛选并结构化原则以形成有效宪法;(2) 评估微调后的CAI模型是否真正遵循这些原则。通过分析来自人工智能与心理学的原则,发现正向表述的行为型原则比负向或特质型原则更符合人类偏好。在安全对齐场景中,采用图结构方法优化现有宪法,提升了安全性能,同时保持了强大的通用推理能力。有趣的是,微调后的模型在负向原则上表现良好,但在正向原则上表现不佳,与人类对齐结果相反,凸显了原则设计与模型实际遵循之间的潜在鸿沟。总体而言,C3AI为宪法型AI的构建与评估提供了结构化、可扩展的方法。
原文摘要 · Abstract (English)
Constitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge. We introduce the C3AI framework (\textit{Crafting Constitutions for CAI models}), which serves two key functions: (1) selecting and structuring principles to form effective constitutions before fine-tuning; and (2) evaluating whether fine-tuned CAI models follow these principles in practice. By analyzing principles from AI and psychology, we found that positively framed, behavior-based principles align more closely with human preferences than negatively framed or trait-based principles. In a safety alignment use case, we applied a graph-based principle selection method to refine an existing CAI constitution, improving safety measures while maintaining strong general reasoning capabilities. Interestingly, fine-tuned CAI models performed well on negatively framed principles but struggled with positively framed ones, in contrast to our human alignment results. This highlights a potential gap between principle design and model adherence. Overall, C3AI provides a structured and scalable approach to both crafting and evaluating CAI constitutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。