arXiv:2412.18497cs.CL2024-12EMNLP被引 2

发现大模型中负责记忆与泛化的神经元,可实时调控其行为。

Neuron-Level Differentiation of Memorization and Generalization in Large Language Models

  • 通过任务设计识别出两类特化神经元
  • 干预神经元能切换模型记忆或泛化模式
  • 结果在不同模型间一致,具备通用性

我们研究大语言模型在神经元层面区分记忆与泛化的行为机制。通过精心设计的任务,识别出分别负责两类行为的神经元子集。在从头训练的GPT-2模型和使用LoRA微调的预训练LLaMA-3.2模型上均观察到一致的神经元特化现象。进一步证明,在推理阶段对这些神经元进行干预,可引导模型行为向记忆或泛化方向转变。通过评估任务内与跨任务的一致性,验证了神经元-行为关联具有普遍性,而非数据集特异性。研究揭示了大模型中的模块化结构,并实现了推理时对记忆与泛化行为的可控调节。

原文摘要 · Abstract (English)

We investigate how Large Language Models (LLMs) distinguish between memorization and generalization at the neuron level. Through carefully designed tasks, we identify distinct neuron subsets responsible for each behavior. Experiments on both a GPT-2 model trained from scratch and a pretrained LLaMA-3.2 model fine-tuned with LoRA show consistent neuron-level specialization. We further demonstrate that inference-time interventions on these neurons can steer the model's behavior toward memorization or generalization. To assess robustness, we evaluate intra-task and inter-task consistency, confirming that these neuron-behavior associations reflect generalizable patterns rather than dataset-specific artifacts. Our findings reveal modular structure in LLMs and enable controlling memorization and generalization behaviors at inference time.

大模型神经元分析记忆控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。