提出验证siRNA预测解释性的新方法,防止错误设计指导。
Validating Interpretability in siRNA Efficacy Prediction: A Perturbation-Based, Dataset-Aware Protocol
- 用反事实扰动测试高重要性位置的修改是否显著改变预测结果。
- 发现20个模型实例中19个通过验证,1个存在反向重要性问题。
- 适合关注药物设计可解释性与可靠性研究者使用。
Saliency maps 在 siRNA 效力预测中日益被用作设计指引,但归因方法在推动序列修改前极少经过验证。本文提出一种合成前验证机制:基于扰动的反事实敏感性一致性测试,检验高重要性位置的突变是否比结构匹配的对照组更显著影响模型输出。跨数据集迁移分析揭示两种此前未被察觉的失败模式:‘忠实但错误’(归因有效但预测失败)和‘反向重要性’(最高重要性编辑的影响小于随机编辑)。令人震惊的是,基于mRNA水平检测训练的模型在荧光素酶报告基因数据集上完全失效,表明协议差异可能无声地使部署无效。在四个基准上,20个交叉验证实例中有19个通过测试;唯一失败案例显示存在反向重要性。引入生物先验正则化(BioPrior)可适度提升归因一致性,带来轻微且依赖数据集的预测性能损失。结果确立了归因验证作为解释引导治疗设计的必要前置流程。代码已公开于 https://github.com/shadi97kh/BioPrior。
原文摘要 · Abstract (English)
Saliency maps are increasingly used as design guidance in siRNA efficacy prediction, yet attribution methods are rarely validated before motivating sequence edits. We introduce a pre-synthesis gate: a protocol for counterfactual sensitivity faithfulness that tests whether mutating high-saliency positions changes model output more than composition-matched controls. Cross-dataset transfer reveals two failure modes that would otherwise go undetected: faithful-but-wrong (saliency valid, predictions fail) and inverted saliency (top-saliency edits less impactful than random). Strikingly, models trained on mRNA-level assays collapse on a luciferase reporter dataset, demonstrating that protocol shifts can silently invalidate deployment. Across four benchmarks, 19/20 fold instances pass; the single failure shows inverted saliency. A biology-informed regularizer (BioPrior) strengthens saliency faithfulness with modest, dataset-dependent predictive trade-offs. Our results establish saliency validation as essential pre-deployment practice for explanation-guided therapeutic design. Code is available at https://github.com/shadi97kh/BioPrior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。