arXiv:2605.02962cs.LGstat.CO2026-05

用干预方法检测药物靶点预测模型是否真懂分子机制。

ISAAC: Auditing Causal Reasoning in Deep Models for Drug-Target Interaction

  • 通过干预输入测试模型对真实与虚假特征的依赖程度。
  • 不同模型间推理得分差达25%,但准确率仅差3%。
  • 适合关注模型可解释性与科学可信度的研究者。

用于药物-靶点相互作用(DTI)预测的深度学习模型常在基准测试中表现优异,却未必依赖于具有生物学意义的分子特征,而传统基于准确率的评估无法发现这一缺陷。本文提出ISAAC(基于干预的结构审计方法),一种后验框架,通过匹配的机制性与伪相关输入干预,独立于预测准确率地评估模型对先验知识的结构敏感性。在Davis基准上应用于三种序列型DTI架构,ISAAC揭示各模型间推理得分存在约25%的相对差异,而其AUROC相近(约3%内),且在不同训练与干预种子、两种扰动算子下均保持稳定。这些差异在常规准确率指标下不可见,表明后验结构审计应作为分子建模中科学机器学习评估的补充手段。

原文摘要 · Abstract (English)

Deep learning models for drug--target interaction (DTI) prediction often achieve strong benchmark performance without necessarily relying on mechanistically meaningful molecular features, a limitation that standard accuracy-based evaluation cannot detect. We introduce ISAAC (Intervention-based Structural Auditing Approach for Causal Reasoning), a post-hoc framework that evaluates prior-relative structural sensitivity by probing frozen models through matched mechanistic and spurious input-level interventions, independently of predictive accuracy. Applied to three sequence-based DTI architectures on the Davis benchmark, ISAAC reveals approximately 25\% relative differences in reasoning scores across models with comparable AUROC (within around 3\%), stable across training and intervention seeds and two distinct perturbation operators. These discrepancies, undetectable under conventional accuracy metrics, motivate the use of post-hoc structural auditing as a complement to standard performance evaluation in scientific machine learning for molecular modeling.

因果推理药物发现模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。