发现语言模型能通过调整隐藏表示实现道德自修正。
Explaining Intrinsic Moral Self-Correction with Mechanistic Interpretability
- 通过可解释的潜在方向控制模型表示,实现道德自修正。
- 在4个任务中验证了提示引发的表示变化与对比向量对齐。
- 激活添加法比原始提示更有效,适合研究模型伦理机制。
内在道德自修正指语言模型仅通过提示就能改进其伦理判断或输出对齐。尽管在多种任务中表现有效,其机制仍不明确。我们假设该现象通过沿可解释的潜在方向引导隐藏表示实现。在六个大模型上评估四个与道德相关的任务,结果表明自修正提示引发的表示变化与对比性引导向量对齐,且这种对齐在使用不相关语料构建的引导向量时仍成立。值得注意的是,通过激活添加方式应用这些提示诱导的偏移,能比自修正提示本身或引导向量更有效地改变模型行为。研究结果表明,表示引导是内在道德自修正的机制核心。
原文摘要 · Abstract (English)
Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely through prompting. While effective across diverse tasks, its mechanism remains unclear. We hypothesize intrinsic moral self-correction functions by steering hidden representations along interpretable latent directions. Evaluating six LLMs across four morality-related tasks, we demonstrate that the representation shifts induced by self-correction prompts align with contrastive steering vectors. This alignment transfers even when the steering vectors are constructed from a disjoint corpus. Notably, when applied via activation addition, these prompt-induced shifts can alter model behavior more effectively than the self-correction prompts and the steering vectors. Our findings suggest representation steering is the mechanistic driver of intrinsic moral self-correction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。