模型准确率高,但解释逻辑可能大不相同,该研究揭示了训练过程中的解释多样性。
EvoXplain: When Machine Learning Models Agree on Predictions but Disagree on Why -- Measuring Mechanistic Multiplicity Across Training Runs
- 将解释视为多次训练产生的样本,分析其是否形成统一解释模式
- 98%准确率下,不同正则化强度导致解释结构分化为多个稳定类别
- 适用于关注模型可解释性稳定性的基因组学与医疗AI研究者
机器学习模型主要依据预测性能评估,尤其在应用基因组学中,解释常被当作生物学发现。实践中,基因面板通过交叉验证、调参网格和重复运行生成的多个模型进行平均、排序或取共识来稳定。这引发一个被忽视的问题:当两个模型达到高准确率时,它们是否依赖相同的内部逻辑?我们提出EvoXplain,一种诊断框架,用于衡量训练和模型选择流程中解释是否唯一确定。不同于分析单一模型,EvoXplain将解释视为从训练流程中抽取的样本,不聚合预测也不构建集成,考察其是否形成单一连贯解释域,或分裂为多个结构化子域。我们在TCGA泛癌队列和乳腺癌亚型任务上评估,使用弹性网逻辑回归和梯度提升树。尽管所有模型准确率约98%,解释结构在不同流程中差异显著。固定数据划分仅改变正则化强度,等效准确率的逻辑回归模型分化为若干离散且可复现的解释基域,跨100次数据划分重现并携带不同生物内涵;而梯度提升流程则收敛至单一基域。同一癌种亚型内,仅通过常规调参即出现此类多态性。EvoXplain使解释结构可视化,揭示平均共识可能并不对应任何单一训练模型,将可解释性重新定义为训练流程的属性而非单个模型的特性。
原文摘要 · Abstract (English)
Machine learning models are primarily judged by predictive performance, especially in applied genomics, where explanations are read as biological findings. In practice, reported gene panels are stabilised by averaging, ranking, or taking consensus over the many models a pipeline produces across cross-validation folds, tuning grids, and repeated runs. This raises an overlooked question: when two models achieve high accuracy, do they rely on the same internal logic, or reach the same outcome via different mechanisms? We introduce EvoXplain, a diagnostic framework that measures whether a pipeline's explanation is uniquely determined across repeated training and model selection. Rather than analysing a single trained model, EvoXplain treats explanations as samples drawn from the training and model selection pipeline itself, without aggregating predictions or constructing ensembles, and examines whether they form a single coherent explanatory basin or separate into multiple structured basins. We evaluate EvoXplain on a TCGA pan-cancer cohort and a within-cancer breast-cancer subtype task, using elastic-net Logistic Regression and gradient-boosted trees. Although all models reach about 98% accuracy, explanation structure differs across pipelines. Holding the data split fixed and varying only the regularisation strength, equally accurate Logistic Regression models separate into a few discrete, reproducible basins that recur across 100 data splits and carry distinct biological content, while the gradient-boosted pipeline converges to one basin. The same multiplicity appears within a single cancer subtype, from the ordinary tuning step alone. EvoXplain makes explanatory structure visible, revealing when an averaged consensus corresponds to no single trained model, and reframes interpretability as a property of the training pipeline rather than of any single model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。