首次实证研究自然科学中预训练模型的复用模式及其影响。
An Empirical Investigation of Pre-Trained Deep Learning Model Reuse in the Scientific Process
- 分析1.77万篇开源论文,量化预训练模型在科研中的使用情况。
- 生物化学领域复用最活跃,模型微调是主要复用方式。
- 模型复用显著提升科研测试阶段效率,适合跨学科研究者参考。
深度学习在自然科学领域已产生重要影响,但从零训练模型的高昂财务与技术成本限制了其应用。借鉴软件工程经验,自然科学家开始复用预训练深度学习模型(PTMs)以分摊成本。尽管已有研究推荐复用模式,但本文首次对自然科学领域中PTM复用模式开展实证研究,基于17,718篇同行评审、开放获取的论文,量化了PTM在科研流程中的使用情况与影响。结果显示,'生物化学、遗传学与分子生物学'领域在PTM复用方面领先于其他自然学科;'适应(adaptation)'是各自然学科中最普遍的复用模式;而科研流程中的'测试(testing)'阶段受到PTM集成的最显著影响。
原文摘要 · Abstract (English)
Deep learning has achieved recognition for its impact within natural sciences, yet the prohibitive financial and technical cost of training models from scratch inhibit adoption. Following software engineering community guidance, natural scientists are reusing pre-trained deep learning models (PTMs) to amortize these costs. While prior works recommend PTM reuse patterns, we present the first empirical study of PTM reuse patterns in the natural sciences, quantifying the utilization and impact of PTM reuse within the scientific process across 17,718 peer reviewed, open access papers. Our results show that "Biochemistry, Genetics and Molecular Biology" has outpaced other natural scientific fields in PTM reuse, "adaptation" reuse is the most prevalent PTM reuse pattern identified across all natural science fields, and the "testing" stage of the scientific process has been most impacted by PTM integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。