GFlowNet生成分子时,表面化学可解释性实则来自模型结构和原子特征,而非训练所得。
Interpreting GFlowNets for Drug Discovery: What probes can and cannot show
- 用梯度显著性和反事实编辑等多方法结合控制实验分析模型可解释性
- 未训练的同构网络在药物相似性上表现接近训练模型(差值<0.005)
- 只有部分子结构特征稳定出现,高分提示不等于学到了化学知识
生成流网络(GFlowNets)通过逐步决策构建分子,但其内部策略不透明,限制了在药物发现中的应用,因化学家需要可解释的结构设计理由。本文对一个合成感知的GFlowNet模型SynFlowNet(以药效评分QED为奖励训练)进行了控制验证的可解释性研究。框架结合梯度显著性、反事实编辑、欠完备因子分析与过完备BatchTopK稀疏自编码器,采用打乱标签控制、RDKit描述符基线、骨架分离数据集、跨种子稳定性及同架构未训练网络等多种对照。结果显示:分子理化性质和功能团可从嵌入中高度解码,但同架构未训练网络的表现几乎与训练模型相当(药物相似性差异<0.005,分子尺寸解码略优)。因此,这种解码能力主要源于图结构和原子特征表示,而非策略训练所获。最终结论更狭窄:过完备自编码器在相近稀疏度下优于匹配的欠完备基线;含卤素、硼等化学富集的子结构检测器在单次运行中出现,但仅少数方向在多种子间稳定;单独置零特征会引发属性特异性解码变化。本研究提供了一套可复用的分子模型可解释性控制协议,明确区分学习到的结构与由架构和输入表示带来的信号。仅凭高探测分数不能作为学习到化学知识的证据。
原文摘要 · Abstract (English)
Generative Flow Networks (GFlowNets) construct molecules through sequential decisions, but their internal policies remain opaque, limiting adoption in drug discovery, where chemists need interpretable rationales for proposed structures. We present a control-validated interpretability study of SynFlowNet, a synthesis-aware GFlowNet trained with a drug-likeness (QED) reward. Our framework combines gradient saliency and counterfactual edits, an undercomplete factor analysis, and an overcomplete BatchTopK sparse autoencoder, evaluated with shuffled-label controls, RDKit-descriptor baselines, a scaffold-disjoint split, cross-seed stability, and an architecture-matched untrained network. These controls materially change the interpretation. Physicochemical properties and functional groups are highly decodable from SynFlowNet embeddings, but an untrained network with the same architecture performs essentially as well as the trained policy (drug-likeness within 0.005 and marginally better molecular-size decoding). Thus, this decodability reflects graph architecture and atom featurization rather than representations acquired through policy training. The surviving conclusions are narrower: the overcomplete autoencoder reconstructs embeddings better than a matched undercomplete baseline at comparable sparsity; chemically enriched substructure detectors, including per-halogen and boron features, emerge in individual runs, although only a small subset of dictionary directions is stable across seeds; and zeroing individual features produces property-specific effects on probe decoding. Beyond SynFlowNet, this study provides a reusable control protocol for molecular-model interpretability, separating learned structure from signals supplied by architecture and input representation. High probe scores alone should not be treated as evidence of learned chemistry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。