基因组模型解释方法需告别个例验证,走向系统化评估
Position: Genomic Model Research Must Move Beyond Anecdotal Evaluation of Interpretability Methods

- 提出分层框架,系统评估基因组可解释性方法
- 实证不同方法对同一预测给出矛盾解释,且无法定位已知调控基序
- 适合关注基因组模型可信度的生物学家与算法研究者
机器学习和算力进步极大提升了人类基因组的预测能力,但生物学家要求模型不仅能预测,还需揭示背后的生物学机制。尽管可解释机器学习(IML)技术被广泛应用,当前研究仍严重依赖个例验证:多数工作仅使用单一IML方法,报告孤立的成功案例。我们在转录因子结合任务上进行基准测试,发现不同IML方法常出现三类问题:(1) 对同一预测给出矛盾解释;(2) 无法准确定位已知调控基序;(3) 无法忠实反映模型内部决策过程。为此,我们主张建立类似临床试验的验证框架,强调设计严谨性和不良反应报告,推动基因组可解释性从‘挑选可信案例’转向对一致性、忠实性和生物学有效性的系统评估。为此,我们提出一个分层评估框架,以指导严谨的评估与报告。
原文摘要 · Abstract (English)
Advances in machine learning and computational power have unlocked the predictive potential of the human genome, yet biologists now demand that these models also elucidate the underlying biological mechanisms. While interpretable machine learning (IML) techniques have been increasingly applied to bridge this gap, there has been a pervasive reliance on anecdotal validation: the vast majority of research relies on a single IML method and reports only isolated successful instances. Through a benchmarking study on transcription factor binding, we demonstrate the risks of current practices. We show that different IML methods can often (1) yield contradictory explanations for the same predictions, (2) fail to localize known regulatory motifs, and (3) fail to faithfully reflect the model's internal decision process. In light of this, we argue for a validation framework analogous to clinical trials: just as trials require rigorous design and adverse-event reporting, genomic interpretability must move beyond cherry-picked plausibility toward systematic assessment of consistency, faithfulness, and biological validity. To facilitate this, we propose a tiered framework to guide rigorous evaluation and reporting of genomic IML methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。