arXiv:2601.09398cs.CLcs.AI2026-01

通过定位关键参数通道,实现大模型能力的精准迁移与恢复

Ability Transfer and Recovery via Modularized Parameters Localization

  • 基于激活差异定位能力相关通道,仅转移特定参数
  • 在多语言数学推理中恢复遗忘能力,干扰极小
  • 适合需要持续学习、避免灾难性遗忘的研究者

大规模语言模型在持续预训练或微调后,虽能提升特定领域、语言或技能表现,但常导致其他能力退化甚至灾难性遗忘。我们通过分析相关模型在领域和语言特定输入下的模块激活,发现能力相关激活集中在少数通道(通常<5%),且这些通道具有良好的解耦性、充分性和稳定性。基于此,提出ACT(激活引导的通道级能力迁移)方法:通过激活差异定位能力相关通道,仅选择性转移对应参数,并进行轻量微调以保证兼容性。多语言数学与科学推理实验表明,ACT可在保留原有技能的同时恢复遗忘能力,还能将多个专用模型合并为单一模型,实现多能力集成且干扰最小。代码与数据将公开发布。

原文摘要 · Abstract (English)

Large language models can be continually pre-trained or fine-tuned to improve performance in specific domains, languages, or skills, but this specialization often degrades other capabilities and may cause catastrophic forgetting. We investigate how abilities are distributed within LLM parameters by analyzing module activations under domain- and language-specific inputs for closely related models. Across layers and modules, we find that ability-related activations are highly concentrated in a small set of channels (typically <5\%), and these channels are largely disentangled with good sufficiency and stability. Building on these observations, we propose ACT (Activation-Guided Channel-wise Ability Transfer), which localizes ability-relevant channels via activation differences and selectively transfers only the corresponding parameters, followed by lightweight fine-tuning for compatibility. Experiments on multilingual mathematical and scientific reasoning show that ACT can recover forgotten abilities while preserving retained skills. It can also merge multiple specialized models to integrate several abilities into a single model with minimal interference. Our code and data will be publicly released.

大模型能力迁移参数定位持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。