arXiv:2410.23496cs.CL2024-10被引 6

小模型经安全对齐后也能自我修正道德问题,突破了规模依赖认知。

Smaller Large Language Models Can Do Moral Self-Correction

  • 通过精心设计提示,3.8B小模型实现良好道德自纠正能力。
  • 小模型理解社会规范与自解释能力弱于大模型,但对错误指令仍无效。
  • 适合关注模型伦理安全、资源受限场景下的研究者参考。

自纠正能力是大型语言模型的新兴特性之一,使模型能在自然语言反馈指导下自我修正不当输出。道德自纠正是一种事后方法,在不需梯度更新的前提下修正非道德生成内容,兼具计算轻量与保留语言建模能力的优势。先前研究表明模型可自我去偏,但普遍认为参数少于220亿的小模型不具备道德自纠正能力。本文通过细致提示实验,验证了该假设在社会刻板印象情境下的有效性。结果表明:(i)出人意料地,经过适当安全对齐微调的3.8B模型可实现优异的道德自纠正表现,凸显安全对齐的关键作用;(ii)小模型在理解社会规范和通过思维链(CoT)进行自解释方面确实弱于大规模模型,但所有规模模型在面对非道德指令时均表现出较差的自纠正性能。

原文摘要 · Abstract (English)

Self-correction is one of the most amazing emerging capabilities of Large Language Models (LLMs), enabling LLMs to self-modify an inappropriate output given a natural language feedback which describes the problems of that output. Moral self-correction is a post-hoc approach correcting unethical generations without requiring a gradient update, making it both computationally lightweight and capable of preserving the language modeling ability. Previous works have shown that LLMs can self-debias, and it has been reported that small models, i.e., those with less than 22B parameters, are not capable of moral self-correction. However, there is no direct proof as to why such smaller models fall short of moral self-correction, though previous research hypothesizes that larger models are skilled in following instructions and understanding abstract social norms. In this paper, we empirically validate this hypothesis in the context of social stereotyping, through meticulous prompting. Our experimental results indicate that (i) surprisingly, 3.8B LLMs with proper safety alignment fine-tuning can achieve very good moral self-correction performance, highlighting the significant effects of safety alignment; and (ii) small LLMs are indeed weaker than larger-scale models in terms of comprehending social norms and self-explanation through CoT, but all scales of LLMs show bad self-correction performance given unethical instructions.

道德对齐小模型自纠正提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。