让闭词汇模型也能用任务相加法高效更新,无需重新训练。
Task Addition and Weight Disentanglement in Closed-Vocabulary Models
- 在闭词汇图像分类模型中验证任务相加法的有效性
- 预训练模型天然具备权重解耦特性,支持任务编辑
- 线性探测已可媲美任务相加,适合资源受限场景
任务算术近年来成为编辑预训练开放词汇模型的有力方法,提供了一种比标准多任务微调更低成本的替代方案。然而,尽管存在大量未使用语言监督进行预训练的闭词汇模型,该方法在这些模型上的应用仍属空白。本文研究了闭词汇图像分类模型中的任务相加法。我们考察了多种预训练方案,发现权重解耦——即支持任务算术的属性——是预训练的普遍结果,出现在不同类型的闭词汇预训练视觉变换器中。事实上,我们发现这些预训练的闭词汇视觉变换器同样可通过任务算术进行编辑,实现高任务相加性能,并支持多任务模型的高效部署。最后,我们证明简单的线性探测是任务相加的有竞争力基线。总体而言,我们的发现将任务算术的应用范围扩展至更广泛的预训练模型类别,为在多样环境中更高效地使用预训练模型开辟了新路径。
原文摘要 · Abstract (English)
Task arithmetic has recently emerged as a promising method for editing pre-trained \textit{open-vocabulary} models, offering a cost-effective alternative to standard multi-task fine-tuning. However, despite the abundance of \textit{closed-vocabulary} models that are not pre-trained with language supervision, applying task arithmetic to these models remains unexplored. In this paper, we deploy and study task addition in closed-vocabulary image classification models. We consider different pre-training schemes and find that \textit{weight disentanglement} -- the property enabling task arithmetic -- is a general consequence of pre-training, as it appears in different pre-trained closed-vocabulary models. In fact, we find that pre-trained closed-vocabulary vision transformers can also be edited with task arithmetic, achieving high task addition performance and enabling the efficient deployment of multi-task models. Finally, we demonstrate that simple linear probing is a competitive baseline to task addition. Overall, our findings expand the applicability of task arithmetic to a broader class of pre-trained models and open the way for more efficient use of pre-trained models in diverse settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。