arXiv:2512.12932cs.LGcs.AI2025-12AAAI被引 1

通过影响力引导剪枝,大幅降低生物模型预训练的数据量与计算成本。

Investigating Data Pruning for Pretraining Biological Foundation Models at Scale

  • 基于样本影响力构建高效剪枝框架,低开销估算序列重要性。
  • 在99%以上极端剪枝率下仍优于随机采样,蛋白与RNA任务均有效。
  • 适合资源有限的实验室,推动可复现、可持续的生物AI研究。

生物基础模型(BioFMs)在大规模生物序列上预训练后,展现出对多种下游生信任务的强大表征能力。然而,这类模型通常依赖数百万至数十亿条训练序列及数十亿参数,导致计算成本极高,阻碍了学术机构的复现与使用。为此,本文研究了生物模型预训练中数据剪枝的可行性,提出一种面向生物领域的后处理影响力引导剪枝框架。该方法引入基于子集的自影响力公式,以低计算代价高效估计样本重要性,并设计了两种简单有效的选择策略:Top-k影响力(Top I)和覆盖导向影响力(CCI)。我们在两个代表性模型RNA-FM和ESM-C上验证了该方法。对于RNA任务,在超过99%的极端剪枝率下,本框架持续优于随机采样基线;在蛋白任务中,其生成的共核集(coreset)表现甚至超越十倍大小的随机子集,揭示了生物序列数据中的显著冗余性。这些结果表明,影响力引导剪枝可显著降低生物基础模型预训练的计算开销,为更高效、可访问、可持续的生物人工智能研究铺平道路。

原文摘要 · Abstract (English)

Biological foundation models (BioFMs), pretrained on large-scale biological sequences, have recently shown strong potential in providing meaningful representations for diverse downstream bioinformatics tasks. However, such models often rely on millions to billions of training sequences and billions of parameters, resulting in prohibitive computational costs and significant barriers to reproducibility and accessibility, particularly for academic labs. To address these challenges, we investigate the feasibility of data pruning for BioFM pretraining and propose a post-hoc influence-guided data pruning framework tailored to biological domains. Our approach introduces a subset-based self-influence formulation that enables efficient estimation of sample importance at low computational cost, and builds upon it two simple yet effective selection strategies, namely Top-k Influence (Top I) and Coverage-Centric Influence (CCI). We empirically validate our method on two representative BioFMs, RNA-FM and ESM-C. For RNA, our framework consistently outperforms random selection baselines under an extreme pruning rate of over 99 percent, demonstrating its effectiveness. Furthermore, we show the generalizability of our framework on protein-related tasks using ESM-C. In particular, our coreset even outperforms random subsets that are ten times larger in both RNA and protein settings, revealing substantial redundancy in biological sequence datasets. These findings underscore the potential of influence-guided data pruning to substantially reduce the computational cost of BioFM pretraining, paving the way for more efficient, accessible, and sustainable biological AI research.

生物AI数据剪枝模型压缩高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。