研究模型在有毒数据微调中的机制退化与修复,发现其具有可逆的神经可塑性。
Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification
- 通过有毒数据微调,观察模型内部机制的局部退化。
- 重训污染模型后,原有效机制可被重建,体现神经可塑性。
- 适用于关注模型鲁棒性与机制可恢复性的研究人员。
先前研究显示,对语言模型进行通用任务微调可增强其底层机制。然而,微调在中毒数据上的影响及其机制变化仍不明确。本研究探究了模型在毒性微调过程中的机制演变,识别出主要的退化机制。我们还分析了将污染模型在原始数据上重新训练后的变化,观察到神经可塑性行为:模型在微调污染模型后能重新学习原始机制。研究发现:(i) 任务特定微调会放大底层机制,且该现象可推广至更长训练周期;(ii) 模型通过毒性微调产生的退化局限于特定电路组件;(iii) 在干净数据上重训污染模型时,模型表现出神经可塑性,能够重构原有机制。
原文摘要 · Abstract (English)
Previous research has shown that fine-tuning language models on general tasks enhance their underlying mechanisms. However, the impact of fine-tuning on poisoned data and the resulting changes in these mechanisms are poorly understood. This study investigates the changes in a model's mechanisms during toxic fine-tuning and identifies the primary corruption mechanisms. We also analyze the changes after retraining a corrupted model on the original dataset and observe neuroplasticity behaviors, where the model relearns original mechanisms after fine-tuning the corrupted model. Our findings indicate that: (i) Underlying mechanisms are amplified across task-specific fine-tuning which can be generalized to longer epochs, (ii) Model corruption via toxic fine-tuning is localized to specific circuit components, (iii) Models exhibit neuroplasticity when retraining corrupted models on clean dataset, reforming the original model mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。