arXiv:2502.11085cs.LGcs.AI2025-02

精选对齐数据预训练,效果远超海量混合数据。

On the Importance of Pretraining Data Alignment for Atomic Property Prediction

  • 用化学相似性指数筛选与任务匹配的预训练数据
  • 仅用1/24预算达成甚至超越大规模联合预训练性能
  • 适合关注数据质量而非数量的研究者

本文挑战了原子性质预测领域普遍认为增大数据量和计算资源即可提升性能的范式。我们发现,基于精心筛选的任务对齐数据集进行预训练,可在仅使用1/24预训练预算的情况下,达到甚至超越大规模联合预训练的效果。为此提出化学相似性指数(CSI),一种受计算机视觉中Fréchet Inception Distance启发的分子图度量,用于量化上游预训练数据集与下游任务之间的对齐程度。通过选择CSI距离最小的数据集,我们证明:在小规模、聚焦的数据集上预训练的模型,在下游任务上的表现始终优于在如JMP等大规模混合数据集上预训练的模型,即使这些混合数据集包含与目标任务最对齐的上游数据。反直觉的是,当新增数据与目标任务对齐性差时,盲目增加数据反而会降低模型性能。研究强调,在原子性质预测的预训练中,质量常胜过数量。

原文摘要 · Abstract (English)

This paper challenges the recent paradigm in atomic property prediction that links progress to growing dataset sizes and computational resources. We show that pretraining on a carefully selected task-aligned dataset can match or even surpass large-scale joint pretraining while using only 1/24th of the pretraining budget. We introduce the Chemical Similarity Index (CSI), a simple metric for molecular graphs inspired by the Fréchet Inception Distance in computer vision, which quantifies the alignment between upstream pretraining datasets and downstream tasks. By selecting the most aligned dataset with minimal CSI distance, we show that models pretrained on a smaller, focused dataset consistently achieve better performance on downstream tasks than those pretrained on massive, mixed datasets such as JMP. This holds even when the mixed dataset includes the upstream dataset most aligned with the downstream task. Counterintuitively, we also find that indiscriminately adding more data can degrade model performance when the additional data is poorly aligned with the target task. Our findings highlight that quality often outperforms quantity in pretraining for atomic property prediction.

原子性质预测数据对齐预训练化学相似性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。