arXiv:2501.08145cs.CLcs.AI2025-01被引 15

发现大模型拒绝行为具有非线性特征,突破传统线性理解

Refusal Behavior in Large Language Models: A Nonlinear Perspective

  • 用降维技术分析六种模型的拒绝行为模式
  • 不同架构和层级的拒绝机制呈现多维非线性特征
  • 为安全对齐研究提供新视角,适合模型可解释性研究者

大型语言模型(LLMs)的拒绝行为使其能够拒绝回应有害、不道德或不当提示,从而确保与伦理标准对齐。本文研究了来自三种架构家族的六种LLM的拒绝行为。通过主成分分析(PCA)、t-SNE和UMAP等降维技术,挑战了将拒绝行为视为线性现象的传统假设。结果表明,拒绝机制具有非线性、多维特性,且因模型架构和层位置而异。这些发现强调了在对齐研究中采用非线性可解释性方法的重要性,并为更安全的AI部署策略提供依据。

原文摘要 · Abstract (English)

Refusal behavior in large language models (LLMs) enables them to decline responding to harmful, unethical, or inappropriate prompts, ensuring alignment with ethical standards. This paper investigates refusal behavior across six LLMs from three architectural families. We challenge the assumption of refusal as a linear phenomenon by employing dimensionality reduction techniques, including PCA, t-SNE, and UMAP. Our results reveal that refusal mechanisms exhibit nonlinear, multidimensional characteristics that vary by model architecture and layer. These findings highlight the need for nonlinear interpretability to improve alignment research and inform safer AI deployment strategies.

拒绝行为非线性可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。