用置信区间提升药物分子属性预测的可靠性,应对真实场景中的数据分布偏移。
Conformal Prediction for Molecular Properties under Label Shift

- 基于标签分布偏移加权校准分数,生成统计严格置信区间。
- 无需重新训练,在属性分布变化时仍保持可靠不确定性估计。
- 适合医药研发中需高可信度决策的场景,助力监管合规。
药物发现与开发支撑医疗健康,但成本高且易失败。关键瓶颈在于预测分子性质(如溶解度、活性、毒性),这些直接影响候选药物能否进入临床试验。人工智能加速了这一过程,但其可靠性常因分布偏移而受损,实验条件与训练数据不一致。此外,传统点预测仅提供单一数值,难以支持高风险实验设计。本文提出针对标签偏移的置信区间预测框架,通过使用边缘标签概率比加权置信评分,无需重训练即可生成统计严谨的预测区间。该方法在属性分布漂移时仍能提供稳健的不确定性量化,直接解决实际药物研发中最普遍的挑战之一。从仅关注准确率转向提供可操作的置信度,显著提升AI预测的可信度,并更契合监管对透明度和不确定性报告的要求,最终支持千亿美元级研发管线的可靠决策。
原文摘要 · Abstract (English)
Drug discovery and development underpins healthcare but remains costly and failure-prone. A critical bottleneck lies in predicting molecular properties such as solubility, potency, and toxicity, which directly determine whether a candidate can advance from preclinical to clinical trials. Artificial Intelligence (AI) has accelerated this process, yet its reliability is often undermined by distribution shift, as experimental conditions frequently diverge from training data. In addition, conventional point predictions provide only single-value estimates, offering limited guidance for high-stakes experimental design. We address these challenges with a conformal prediction framework tailored to label shift. By weighting conformal scores using marginal label probability ratios, our method produces statistically rigorous prediction intervals without retraining. This enables robust uncertainty quantification even when property distributions drift, directly tackling one of the most pervasive obstacles to applying AI in real-world drug development. By moving beyond accuracy alone to provide actionable confidence measures, our approach enhances the trustworthiness of AI-driven predictions. This further aligns predictive modeling with regulatory demands for transparency and uncertainty reporting and ultimately supports more reliable decision-making in billion-dollar development pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。