arXiv:2605.26183q-bio.QMcs.LG2026-05被引 1

图神经网络无法解释药物45%的副作用,研究提出四类可解释性缺口。

What Molecular Structure Cannot Tell Us: A Taxonomy of Explainability Gaps in GNN-Based Drug Toxicity Prediction

  • 基于分子结构构建模型时,仅能推断45%已知副作用
  • 发现四种结构性信息缺失类型,包括数据缺失与表征误差
  • 对药物安全评估和监管框架有直接指导意义

并非所有临床相关的不良反应都能从分子图中推断——无论模型质量或架构复杂度如何。本研究提出一种操作性分类体系,揭示了基于结构的毒性预测在信息层面的固有限制,且独立于学习算法。以阿司匹林(ASA)为典型药物,使用Tox21基准训练消息传递神经网络(MPNN),并用GNNExplainer进行原子级归因分析。结果显示,分子结构仅能解释阿司匹林11种已知不良反应中的5种(约45%)。提出四类缺口分类(GAP-1至GAP-4),涵盖原则上不可编码效应、由缺失非随机(MNAR)机制引发的数据缺口、检测方法不匹配及表征误差。通过系统查询ChEMBL数据库,实证量化出42个已记录实验中无有效生物活性数据。注意力池化实验表明表征误差源于MPNN的消息传递层而非聚合步骤。该分类体系对药物安全信号检测与监管规范(如良好药物警戒实践GVP和新方法学NAMs)具重要意义。在药物相互作用(DDI)消融研究中进一步验证了这些结构限制。

原文摘要 · Abstract (English)

Not all clinically relevant adverse effects are structurally inferable from molecular graphs - regardless of model quality or architectural complexity. This study introduces an operational taxonomy of the structural information limits that prevent structure-based toxicity prediction, independent of the learning algorithm employed. Graph Neural Networks (GNNs) have emerged as a natural approach for molecular toxicity prediction, operating directly on atomic connectivity without the information loss inherent to fixed-length fingerprints. However, the fraction of a drug's known pharmacological profile that is actually inferable from molecular structure remains systematically underexplored. A systematic case study using acetylsalicylic acid (ASA, Aspirin) - one of the most comprehensively characterized drugs in pharmacology - serves as model compound. A Message Passing Neural Network (MPNN) is trained on the Tox21 benchmark and GNNExplainer is applied to characterize atom-level attribution. Results indicate that molecular structure explains approximately 45% (5/11) of known ASA adverse effects. A four-category Gap Taxonomy (GAP-1 through GAP-4) is introduced distinguishing between principally non-encodable effects, data gaps arising from Missing Not At Random (MNAR) mechanisms, assay panel mismatches, and representation errors. The MNAR gap is empirically quantified via a systematic ChEMBL query (42 documented assays, 0 retrievable bioactivity entries). An attention pooling experiment localizes the representation error to the MPNN message passing layers rather than the aggregation step. The Gap Taxonomy has direct implications for drug safety signal detection and regulatory frameworks including Good Pharmacovigilance Practice (GVP) guidelines and New Approach Methodologies (NAMs). Structural limits identified are confirmed in a companion DDI ablation study.

图神经网络药物毒性可解释性数据缺口

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。