解决药物发现中数据选择偏差问题,提升小样本分子属性预测效果
Contextual Representation Anchor Network to Alleviate Selection Bias in Few-Shot Drug Discovery
- 用分子表示的聚类中心作锚点,融合上下文知识增强表示能力
- 在MoleculeNet和FS-Mol上分别提升2.60%和3.28%的AUC与ΔAUC-PR
- 适合小样本药物发现、需处理非随机实验数据的研究者使用
药物发现中候选药物筛选成功率低,常导致标签数据不足,引发分子属性预测的小样本学习问题。现有方法忽视了化学实验中非随机采样带来的样本选择偏差,影响数据代表性并降低性能。为此,提出上下文表示锚网络(CRA),其中锚点为分子表示的聚类中心,作为桥梁将任务相关上下文知识注入分子表示以增强其表达力。CRA引入双增强机制:上下文增强动态检索相似未标注分子,捕获其任务特定上下文知识以优化锚点;锚点增强利用锚点扩充分子表示。在MoleculeNet和FS-Mol基准及领域迁移实验中评估,CRA在AUC和ΔAUC-PR指标上分别优于当前最优方法2.60%和3.28%,展现出更强泛化能力。
原文摘要 · Abstract (English)
In the drug discovery process, the low success rate of drug candidate screening often leads to insufficient labeled data, causing the few-shot learning problem in molecular property prediction. Existing methods for few-shot molecular property prediction overlook the sample selection bias, which arises from non-random sample selection in chemical experiments. This bias in data representativeness leads to suboptimal performance. To overcome this challenge, we present a novel method named contextual representation anchor Network (CRA), where an anchor refers to a cluster center of the representations of molecules and serves as a bridge to transfer enriched contextual knowledge into molecular representations and enhance their expressiveness. CRA introduces a dual-augmentation mechanism that includes context augmentation, which dynamically retrieves analogous unlabeled molecules and captures their task-specific contextual knowledge to enhance the anchors, and anchor augmentation, which leverages the anchors to augment the molecular representations. We evaluate our approach on the MoleculeNet and FS-Mol benchmarks, as well as in domain transfer experiments. The results demonstrate that CRA outperforms the state-of-the-art by 2.60% and 3.28% in AUC and $Δ$AUC-PR metrics, respectively, and exhibits superior generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。