用少量新数据优化材料模型嵌入,提升所有下游任务性能。
Refining embeddings with fill-tuning: data-efficient generalised performance improvements for materials foundation models
- 通过粗糙度分析定位嵌入空间缺陷区,生成针对性补全数据。
- 仅添加100条数据,使多个任务性能提升近1%。
- 适合希望低成本提升通用模型性能的研究者。
预训练基础模型学习的嵌入可应用于多种下游任务,但若在特定任务上表现不佳,通常需微调以提升性能。然而现有方法必然导致分布外任务性能下降。本文提出'fill-tuning'新方法,通过生成不适用于特定下游任务但能修正嵌入空间缺陷的数据集,实现基础模型的持续预训练。利用粗糙度分析揭示潜在空间拓扑结构,指导生成最具价值的数据。在基于约10^9条数据训练的多款先进材料基础模型上应用fill-tuning,仅增加100条数据即实现所有下游任务接近1%的性能提升。该方法以微调级别的计算成本,实现了基础模型的普遍性能改进。
原文摘要 · Abstract (English)
Pretrained foundation models learn embeddings that can be used for a wide range of downstream tasks. These embeddings optimise general performance, and if insufficiently accurate at a specific task the model can be fine-tuned to improve performance. For all current methodologies this operation necessarily degrades performance on all out-of-distribution tasks. In this work we present 'fill-tuning', a novel methodology to generate datasets for continued pretraining of foundation models that are not suited to a particular downstream task, but instead aim to correct poor regions of the embedding. We present the application of roughness analysis to latent space topologies and illustrate how it can be used to propose data that will be most valuable to improving the embedding. We apply fill-tuning to a set of state-of-the-art materials foundation models trained on $O(10^9)$ data points and show model improvement of almost 1% in all downstream tasks with the addition of only 100 data points. This method provides a route to the general improvement of foundation models at the computational cost of fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。