用数学距离检测科学AI模型是否真懂结构,避免误判预测能力。
I-SAFE: Wasserstein Coherence Metrics for Structural Auditing of Scientific AI Models
- 通过输入扰动评估模型输出分布与领域先验的一致性。
- 在丹尼斯激酶数据集上发现三模型响应差异显著但准确率接近。
- 适合关注模型可解释性与科学可信度的研究者使用。
深度学习模型在科学预测任务中广泛应用,但其高性能常被误认为具备科学意义。实际上,模型可能依赖捷径特征、数据特定规律或分布偏差,在测试集上表现良好却与领域结构无关。为此,我们提出I-SAFE框架,基于Wasserstein一致性度量(WCM)对科学AI模型进行后验分布审计。该框架利用外部结构先验,对黑箱模型施加结构引导的输入扰动,通过三个互补指标评估输出分布一致性:分位数度量(QBM)用于位置一致性,WCM用于序次一致性,以及一种平移不变变体用于形状一致性。我们在药物-靶点相互作用(DTI)预测任务中应用I-SAFE,采用Davis激酶基准数据集、KLIFS结合口袋注释及三种序列基DTI模型(DeepConvDTI、DeepDTA、TAPB)。尽管三模型预测性能相近,I-SAFE揭示出其分布响应模式存在显著差异,这一区别无法通过准确率识别。该框架不依赖具体模型,适用于任何可结构分解输入且有外部先验信息的领域。
原文摘要 · Abstract (English)
Deep learning models are increasingly used in scientific prediction tasks where strong benchmark performance is often interpreted as evidence of scientifically meaningful behavior. This interpretation is fragile, as models may exploit shortcut features, dataset-specific regularities, or distributional biases that are predictive on held-out data but not aligned with domain-relevant structure. To address this limitation, we introduce the \textsc{I-SAFE} (Interventional Secure, Accurate, Fair and Explainable) framework, a post-hoc distributional auditing framework for scientific AI models centered on the Wasserstein Coherence Metric (WCM). Given a trained black-box predictor and an external structural prior encoding domain knowledge about task-relevant input structure, \textsc{I-SAFE} evaluates raw model outputs under structurally guided perturbations of the input. The proposed audit measures output-distribution coherence through three complementary metrics: a Quantile-Based Metric (QBM) for location-level coherence, the WCM for ordinal coherence, and a translation-invariant WCM variant for shape coherence. We instantiate \textsc{I-SAFE} on drug--target interaction (DTI) prediction using the Davis kinase benchmark, KLIFS (Kinase--Ligand Interaction Fingerprints and Structures) binding-pocket annotations, and three sequence-based DTI models: DeepConvDTI, DeepDTA, and TAPB. Although the models operate in a comparable predictive regime, \textsc{I-SAFE} reveals substantially different distributional response profiles, a distinction invisible to accuracy-based evaluation. The framework is model-agnostic and applicable to any domain where inputs admit a structured decomposition and an external prior is available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。