arXiv:2409.01081cs.LGcs.AI2024-09NeurIPS被引 13

用新方法剪掉70%分子数据,还能比全量训练效果更好

Beyond Efficiency: Molecular Data Pruning for Enhanced Generalization

  • 基于损失差异设计评分函数,动态判断样本重要性
  • 在HIV和PCBA数据集上剪掉60%-70%数据仍超全量训练表现
  • 无需源数据即可适配预训练模型,适合迁移学习场景

随着分子任务多样化和大规模数据集的出现,如何高效训练成为亟待解决但研究不足的问题。数据剪枝(DP)通过筛选低影响力样本形成子集以减轻训练负担,但现有方法难以适配依赖预训练模型的分子任务。为此,本文提出面向泛化增强的分子数据剪枝框架MolPeg,聚焦无源数据剪枝场景——即在使用预训练模型时进行剪枝。通过维持两个更新速度不同的模型,引入基于损失差异的新评分函数衡量样本信息量。作为即插即用框架,MolPeg能同时感知源域与目标域,在四个下游任务中持续优于现有剪枝方法。尤为显著的是,在HIV和PCBA数据集上,即使剪除60%-70%数据,其性能仍超越全量数据训练结果。本工作表明,有效的数据剪枝度量可为迁移学习中的效率与泛化提升提供可行路径。

原文摘要 · Abstract (English)

With the emergence of various molecular tasks and massive datasets, how to perform efficient training has become an urgent yet under-explored issue in the area. Data pruning (DP), as an oft-stated approach to saving training burdens, filters out less influential samples to form a coreset for training. However, the increasing reliance on pretrained models for molecular tasks renders traditional in-domain DP methods incompatible. Therefore, we propose a Molecular data Pruning framework for enhanced Generalization (MolPeg), which focuses on the source-free data pruning scenario, where data pruning is applied with pretrained models. By maintaining two models with different updating paces during training, we introduce a novel scoring function to measure the informativeness of samples based on the loss discrepancy. As a plug-and-play framework, MolPeg realizes the perception of both source and target domain and consistently outperforms existing DP methods across four downstream tasks. Remarkably, it can surpass the performance obtained from full-dataset training, even when pruning up to 60-70% of the data on HIV and PCBA dataset. Our work suggests that the discovery of effective data-pruning metrics could provide a viable path to both enhanced efficiency and superior generalization in transfer learning.

分子生成数据剪枝迁移学习预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。