arXiv:2607.06605cs.LGstat.ML2026-07

主流预测方法在药物筛选中忽视稀有靶点,新方法可精准修复。

A Quiet Failure in Calibrated Virtual Screening: Marginal Conformal Prediction Under-Covers the Minority Class, and a Class-Conditional Fix Recovers It

论文配图:A Quiet Failure in Calibrated Virtual Screening: Marginal Conformal Prediction Under-Covers the Minority Class, and a Class-Conditional Fix Recovers It
图 1 · 摘自论文原文
  • 按类别分别校准预测,避免少数类被忽略
  • 稀有类别覆盖率从4.2%提升至目标90%
  • 适合药物研发中需关注罕见毒性或活性的场景

共价预测被用于药物发现以真实评估模型可靠性:设定误差率α,该方法返回的预测集包含真实标签的概率不低于1−α。我们发现,在数据不平衡情况下,这一保证可能具有危险性。在四个数据集上,标准(边际)共价预测虽达到全局90%覆盖目标,但少数类严重暴露:血脑屏障穿透性中实际覆盖率仅64.8%,临床试验毒性中低至4.2%,稀有类别几乎被放弃。此问题不依赖特定模型——随机森林、图神经网络和冻结的化学语言模型均重现该现象(每项p < 0.001),严重程度与稀有标签的基础校准水平相关,而非模型结构。一个守恒关系解释了该效应:少数类覆盖率不足等于多数类盈余乘以不平衡比,可预测实测差距,误差不超过1个百分点,并准确排序各数据集的严重性。该缺陷在合理的骨架划分和另一共价评分下仍存在,而整体准确率和总体覆盖率仍看似良好,因此极易被忽视。类别条件(Mondrian)共价预测在所有数据集上填补了差距,将少数类覆盖率恢复至目标值,仅小幅增加预测集大小。我们定位失败源于通用分子骨架——苯和吡啶核心同时存在于两类中——提出单数值诊断方法,并通过成本模型证明:对受影响化合物选择不预测,可使筛选项目从净负收益转为净正收益。本研究揭示了在真实化学数据中,共价理论中的已知缺口如何在不平衡下变得严重且隐蔽,并提供了一套实用方案以恢复各类别的可靠性。

原文摘要 · Abstract (English)

Conformal prediction is being adopted in drug discovery to put an honest number on model reliability: pick an error rate alpha, and the method returns prediction sets containing the true label with probability at least 1 - alpha. We show this guarantee can be dangerous on imbalanced datasets. Across four datasets, standard (marginal) conformal prediction hits its global 90% coverage target while leaving the minority class badly exposed: realized minority coverage falls to 64.8% on blood-brain-barrier penetration and to 4.2% on clinical-trial toxicity, where the rare class is nearly abandoned. The failure is not tied to one model: a random forest, a graph network, and a frozen chemical language model all reproduce it (p < 0.001 in every case), with severity tracking baseline calibration on rare labels rather than architecture. A conservation identity explains the effect: the minority's shortfall equals the majority's surplus amplified by the imbalance ratio, predicting the measured gap to within one point and ordering severity across datasets. The failure survives realistic scaffold splits and a second conformal score, while aggregate accuracy and overall coverage stay reassuringly high, which is exactly why it is easy to miss. Class-conditional (Mondrian) conformal prediction closes the gap on every dataset, restoring minority coverage to target for a modest increase in prediction-set size. We localize the failures to generic molecular scaffolds - plain benzene and pyridine cores occurring in both classes - propose a one-number diagnostic, and show with a cost model that abstaining on affected compounds flips a screening campaign from net-negative to net-positive utility. Our contribution is demonstrating on real chemistry how severe and invisible this known conformal-theory gap becomes under imbalance, and laying out a practical protocol restoring per-class reliability.

药物发现共价预测不平衡数据分类校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。