通过逐层剥离上下文,提升分子属性预测的准确与可解释性
Peeling Context from Cause for Molecular Property Prediction
- 逐层分离因果信号与非因果上下文,融合多图表示
- 在四个基准上均优于基线,降低误差并提升 $R^2$
- 生成原子级因果显著图,助力分子结构优化设计
深度模型广泛用于分子属性预测,但常因难以解释且依赖虚假上下文而非因果结构,导致分布外性能下降。本文提出CLaP(Causal Layerwise Peeling)框架,通过逐层软分割将因果与非因果分支分离,融合多模态因果证据,并逐步去除批次相关上下文,聚焦于标签相关的结构信息,从而抑制捷径信号并稳定层间优化。在四个分子基准测试中,CLaP consistently 在 MAE、MSE 与 $R^2$ 上超越竞争模型。该模型还能生成原子级因果显著图,突出决定预测的关键子结构,提供可操作的分子改造指引。案例研究验证了显著图的准确性及其与化学直觉的一致性。通过逐层剥离上下文、保留因果信号,模型实现兼具高精度与可解释性的分子设计预测。
原文摘要 · Abstract (English)
Deep models are used for molecular property prediction, yet they are often difficult to interpret and may rely on spurious context rather than causal structure, which reduces reliability under distribution shift and harms predictive performance. We introduce CLaP (Causal Layerwise Peeling), a framework that separates causal signal from context in a layerwise manner and integrates diverse graph representations of molecules. At each layer, a causal block performs a soft split into causal and non-causal branches, fuses causal evidence across modalities, and progressively removes batch-coupled context to focus on label-relevant structure, thereby limiting shortcut signals and stabilizing layerwise refinement. Across four molecular benchmarks, CLaP consistently improves MAE, MSE, and $R^2$ over competitive baselines. The model also produces atom-level causal saliency maps that highlight substructures responsible for predictions, providing actionable guidance for targeted molecular edits. Case studies confirm the accuracy of these maps and their alignment with chemical intuition. By peeling context from cause at every layer, the model yields predictors that are both accurate and interpretable for molecular design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。