arXiv:2506.11877cs.LGcs.AI2025-06被引 2

用无标签数据插值,提升分子属性预测在分布外化合物上的泛化能力。

Robust Molecular Property Prediction via Densifying Scarce Labeled Data

  • 通过双层优化,利用无标签数据在分布内与分布外样本间插值
  • 在存在显著协变量偏移的真实数据集上实现显著性能提升
  • 特别适合标签稀缺且需预测未知化合物的药物研发场景

分子属性预测模型普遍依赖训练数据中出现的结构,导致对分布外化合物的泛化能力差。在药物发现中,最关键的化合物往往超出训练集范围,这种偏差尤为严重。该不匹配引发显著协变量偏移,使标准深度学习模型预测不稳定且不准确。此外,因实验验证成本高昂,标注数据稀少,进一步加剧了可靠泛化的难度。为此,我们提出一种新颖的双层优化方法,利用无标签数据在分布内(ID)与分布外(OOD)数据间进行插值,使模型学会超越训练分布的泛化能力。我们在多个具有显著协变量偏移的现实世界数据集上展示了显著性能提升,并通过 t-SNE 可视化验证了插值方法的有效性。

原文摘要 · Abstract (English)

A widely recognized limitation of molecular prediction models is their reliance on structures observed in the training data, resulting in poor generalization to out-of-distribution compounds. Yet in drug discovery, the compounds most critical for advancing research often lie beyond the training set, making the bias toward the training data particularly problematic. This mismatch introduces substantial covariate shift, under which standard deep learning models produce unstable and inaccurate predictions. Furthermore, the scarcity of labeled data-stemming from the onerous and costly nature of experimental validation-further exacerbates the difficulty of achieving reliable generalization. To address these limitations, we propose a novel bilevel optimization approach that leverages unlabeled data to interpolate between in-distribution (ID) and out-of-distribution (OOD) data, enabling the model to learn how to generalize beyond the training distribution. We demonstrate significant performance gains on challenging real-world datasets with substantial covariate shift, supported by t-SNE visualizations highlighting our interpolation method.

分子预测分布外数据稀缺插值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。