用可组合的软标记实现模型技能持续学习,不更新权重
Skill Neologisms: Towards Skill-based Continual Learning

- 引入技能新词(软标记)替代参数更新
- 合成任务中新词可提升特定技能表现
- 零样本组合已学技能,适合动态扩展能力
现代大语言模型在不断增长的技能上表现出色,并能灵活组合。但以可扩展方式扩展新技能仍是开放问题:微调及其高效变体易导致灾难性遗忘,基于上下文的方法表达能力有限且受限于模型有效上下文长度。本文探索‘技能新词’——即嵌入模型词汇表的软标记,通过优化以增强特定技能能力,而无需更新权重。我们首先发现预训练模型已存在与程序性知识相关的标记。随后在可控的合成任务中表明,技能新词可被学习以提升特定技能表现,且能与分布外技能组合,独立训练的新词可实现零样本组合。最后,在更真实的自然语言任务——Skill-Mix基准上验证了独立学习技能新词的零样本组合能力。结果表明,技能新词可能为基于技能的持续学习提供可扩展路径。
原文摘要 · Abstract (English)
Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending model capabilities to new skills in a scalable manner is an open problem: fine-tuning and parameter-efficient variants risk catastrophic forgetting, while context-based approaches have limited expressiveness and are constrained by the model's effective context. We explore skill neologisms--soft tokens integrated in the model's vocabulary and optimized to improve capabilities over a specific skill--as a way to selectively acquire new skills without weight updates. We first observe that pretrained LLMs already exhibit tokens associated with procedural knowledge. We then show on a controlled synthetic task that skill neologisms can be learned to improve model capabilities on specific skills while being composable with out-of-distribution skills, and that independently trained skill neologisms can be composed zero-shot. Finally, we validate zero-shot composition of independently learned skill neologisms on the more realistic natural language setting of the Skill-Mix benchmark. These results suggest that skill neologisms may provide a scalable path towards skill-based continual learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。