小样本下用自纠正蒸馏提升模型领域适应能力,防止过拟合导致泛化下降
Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation
- 通过样本级自纠正蒸馏实现低数据域适应
- 在仅500样本时仍保持良好泛化,性能优于标准微调2-10倍
- 可与参数高效微调结合,适合资源受限场景
在新领域微调语言模型时,通用性能常会退化,且数据越少问题越严重。本文提出最小微调(MFT),一种无需重用预训练数据即可缓解过拟合导致的泛化下降的方法。MFT在多种模型和领域中表现出2-10倍更优的专化-泛化比,在新域数据少至500样本时仍具内在抗过拟合能力。其采用个体化样本级的自纠正蒸馏机制,性能优于参数高效微调方法,具备类似回放的泛化保护效果,且可与之组合使用以获得协同增益。
原文摘要 · Abstract (English)
Finetuning language models for a new domain inevitably leads to the deterioration of their general performance. This becomes more pronounced the more limited the finetuning data resource. We introduce minifinetuning (MFT), a method for language model domain adaptation that considerably reduces the effects of overfitting-induced degeneralization in low-data settings and which does so in the absence of any pre-training data for replay. MFT demonstrates 2-10x more favourable specialization-to-degeneralization ratios than standard finetuning across a wide range of models and domains and exhibits an intrinsic robustness to overfitting when data in the new domain is scarce and down to as little as 500 samples. Employing corrective self-distillation that is individualized on the sample level, MFT outperforms parameter-efficient finetuning methods, demonstrates replay-like degeneralization mitigation properties, and is composable with either for a combined effect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。