arXiv:2601.11135cs.LGcs.AI2026-01被引 1

通过因果推理提升少样本分子属性预测准确率

Context-aware Graph Causality Inference for Few-Shot Molecular Property Prediction

  • 构建上下文图融合官能团、分子与属性的化学先验知识
  • 在少样本场景下实现更高精度和样本效率
  • 可解释性强,发现的结构符合化学常识

分子属性预测正成为图学习在网页服务中的重要应用,如在线蛋白质结构预测和药物发现。少样本场景下仅少量标记分子可用,是主要挑战。现有研究虽采用上下文学习捕捉分子与属性间关系,但存在两点局限:(1)未利用与属性有因果关联的官能团先验知识;(2)难以识别直接相关的关键子结构。本文提出CaMol框架,从因果推理视角出发,假设每个分子包含决定特定属性的潜在因果结构。首先,引入上下文图,通过连接官能团、分子和属性编码化学知识,引导因果子结构发现;其次,提出可学习的原子掩码策略,分离因果子结构与混淆因素;第三,设计分布干预器,结合因果子结构与化学基础混淆因子,实施后门调整,剥离真实化学变异对因果效应的影响。在多个分子数据集上的实验表明,CaMol在少样本任务中达到更优准确率和样本效率,且对未见属性具有强泛化能力。此外,发现的因果子结构与官能团化学知识高度一致,验证了模型可解释性。

原文摘要 · Abstract (English)

Molecular property prediction is becoming one of the major applications of graph learning in Web-based services, e.g., online protein structure prediction and drug discovery. A key challenge arises in few-shot scenarios, where only a few labeled molecules are available for predicting unseen properties. Recently, several studies have used in-context learning to capture relationships among molecules and properties, but they face two limitations in: (1) exploiting prior knowledge of functional groups that are causally linked to properties and (2) identifying key substructures directly correlated with properties. We propose CaMol, a context-aware graph causality inference framework, to address these challenges by using a causal inference perspective, assuming that each molecule consists of a latent causal structure that determines a specific property. First, we introduce a context graph that encodes chemical knowledge by linking functional groups, molecules, and properties to guide the discovery of causal substructures. Second, we propose a learnable atom masking strategy to disentangle causal substructures from confounding ones. Third, we introduce a distribution intervener that applies backdoor adjustment by combining causal substructures with chemically grounded confounders, disentangling causal effects from real-world chemical variations. Experiments on diverse molecular datasets showed that CaMol achieved superior accuracy and sample efficiency in few-shot tasks, showing its generalizability to unseen properties. Also, the discovered causal substructures were strongly aligned with chemical knowledge about functional groups, supporting the model interpretability.

分子属性预测因果推理少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。