在基因组学等专业领域,简单监督模型仍优于大型预训练模型。
Specialized Foundation Models Struggle to Beat Supervised Baselines
- 用目标数据直接训练简单模型,无需大规模预训练
- 轻量级ResNet或UNet在三个领域均超越最新基础模型
- 适合想快速验证模型性能的研究者和工程团队
继视觉与文本领域成功后,基础模型(FM)范式——即在海量数据上预训练大模型,再针对目标任务微调——已迅速扩展至基因组学、遥感影像和时间序列等科学与工程领域。这一进展是否实现了原始基础模型的突破性效果,即取代传统监督学习?为此,我们考察了三个模态:基因组学、卫星成像与时间序列,并对比了多个近期基础模型与标准监督学习流程的表现:仅使用目标任务数据进行模型开发、超参数调优与训练。结果显示,在这三个专业领域中,均可通过训练简单监督模型(如轻度修改的宽ResNet或UNet)达到甚至超过最新基础模型的性能。本研究表明,大规模预训练的优势在诸多专业领域尚未充分体现,强调应将新基础模型与强且经过良好调优的基线进行比较,并引入两个开源、自动化、易用的新工作流以支持此类评估。
原文摘要 · Abstract (English)
Following its success for vision and text, the "foundation model" (FM) paradigm -- pretraining large models on massive data, then fine-tuning on target tasks -- has rapidly expanded to domains in the sciences, engineering, healthcare, and beyond. Has this achieved what the original FMs accomplished, i.e. the supplanting of traditional supervised learning in their domains? To answer we look at three modalities -- genomics, satellite imaging, and time series -- with multiple recent FMs and compare them to a standard supervised learning workflow: model development, hyperparameter tuning, and training, all using only data from the target task. Across these three specialized domains, we find that it is consistently possible to train simple supervised models -- no more complicated than a lightly modified wide ResNet or UNet -- that match or even outperform the latest foundation models. Our work demonstrates that the benefits of large-scale pretraining have yet to be realized in many specialized areas, reinforces the need to compare new FMs to strong, well-tuned baselines, and introduces two new, easy-to-use, open-source, and automated workflows for doing so.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。