arXiv:2502.04602cs.CLcs.AI2025-02NAACL被引 7

发现大模型对齐中大量依赖浅层知识,可被高效提取复用。

Extracting and Understanding the Superficial Knowledge in Alignment

  • 提出方法分离并提取模型对齐中的浅层知识,仅改最终词元选择
  • 发现安全类任务中浅层知识占比超半数,但推理仍需深层理解
  • 浅层知识可迁移复用,适合快速对齐大模型或修复受损模型

大语言模型与人类价值观对齐通常依赖基于人类反馈的微调,但成本高昂。近期研究发现,通过上下文学习等简单方法也可实现对齐,引发疑问:对齐是否主要依赖浅层知识?本文量化分析该问题,将浅层知识定义为仅通过词元重排即可获得、不改变深层因果关系的知识。提出方法从对齐模型中提取并隔离浅层知识,聚焦于最终词元选择过程的浅层修改。对比仅含浅层知识的模型与完整对齐模型,量化得出:在安全和去毒任务中,浅层知识占主导地位;但涉及推理与上下文理解的任务仍依赖深层知识。此外,实验表明,提取的浅层知识具备可迁移性,可用于高效对齐更大模型;且可恢复,可在模型受损时重建对齐而不影响性能。

原文摘要 · Abstract (English)

Alignment of large language models (LLMs) with human values and preferences, often achieved through fine-tuning based on human feedback, is essential for ensuring safe and responsible AI behaviors. However, the process typically requires substantial data and computation resources. Recent studies have revealed that alignment might be attainable at lower costs through simpler methods, such as in-context learning. This leads to the question: Is alignment predominantly superficial? In this paper, we delve into this question and provide a quantitative analysis. We formalize the concept of superficial knowledge, defining it as knowledge that can be acquired through easily token restyling, without affecting the model's ability to capture underlying causal relationships between tokens. We propose a method to extract and isolate superficial knowledge from aligned models, focusing on the shallow modifications to the final token selection process. By comparing models augmented only with superficial knowledge to fully aligned models, we quantify the superficial portion of alignment. Our findings reveal that while superficial knowledge constitutes a significant portion of alignment, particularly in safety and detoxification tasks, it is not the whole story. Tasks requiring reasoning and contextual understanding still rely on deeper knowledge. Additionally, we demonstrate two practical advantages of isolated superficial knowledge: (1) it can be transferred between models, enabling efficient offsite alignment of larger models using extracted superficial knowledge from smaller models, and (2) it is recoverable, allowing for the restoration of alignment in compromised models without sacrificing performance.

模型对齐浅层知识知识提取可迁移性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。