arXiv:2503.17365cs.LGcs.AI2025-03被引 5

测试小模型用宪法式AI是否有效防有害内容

How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers

  • 用自批判机制让小模型自我纠错
  • 基于Llama的模型减少有害内容明显
  • 不同架构效果差异大,适合评估对齐方法

近期事件凸显大型语言模型(LLMs)的安全风险,推动了对齐方法如宪法式AI(CAI)的研究。本文研究了在小型无审查的7-9B参数模型(DeepSeek-R1-8B、Gemma-2-9B、Llama 3.1-8B、Qwen2.5-7B)上应用CAI自批判机制的效果。结果显示,基于Llama的模型通过自批判显著降低了有害内容生成,而其他架构在剔除后对有害内容检测的改进较小。这表明CAI的有效性可能依赖于模型架构与推理能力。

原文摘要 · Abstract (English)

Recent incidents highlight safety risks in Large Language Models (LLMs), motivating research into alignment methods like Constitutional AI (CAI). This paper explores CAI's self-critique mechanism on small, uncensored 7-9B parameter models: DeepSeek-R1-8B, Gemma-2-9B, Llama 3.1-8B, and Qwen2.5-7B. We show that while Llama-based models exhibited significant harm reduction through self-critique, other architectures demonstrated less improvement in harm detection after abliteration. These results suggest CAI's effectiveness may vary depending on model architecture and reasoning capabilities.

宪法式AI小模型对齐安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。