arXiv:2409.04313cs.LG2024-09被引 8

用生存分析方法提升药物发现中的不确定性估计

Enhancing Uncertainty Quantification in Drug Discovery with Censored Regression Labels

  • 引入Tobit模型处理实验数据中的截断标签
  • 在稀疏数据下显著提升预测不确定性准确性
  • 适合药物研发中资源有限场景的决策支持

药物发现早期阶段的实验决策常依赖计算模型,但实验耗时且成本高昂,因此准确量化机器学习预测的不确定性至关重要。现有方法在数据稀疏、观测稀少的情况下难以有效利用包含阈值信息的截断标签(censored labels)。本文将集成学习、贝叶斯与高斯模型结合生存分析中的Tobit模型,实现对截断标签的学习。结果表明,尽管截断标签信息不完整,仍能显著提升对真实药物研发环境的建模准确性和可靠性。

原文摘要 · Abstract (English)

In the early stages of drug discovery, decisions regarding which experiments to pursue can be influenced by computational models. These decisions are critical due to the time-consuming and expensive nature of the experiments. Therefore, it is becoming essential to accurately quantify the uncertainty in machine learning predictions, such that resources can be used optimally and trust in the models improves. While computational methods for drug discovery often suffer from limited data and sparse experimental observations, additional information can exist in the form of censored labels that provide thresholds rather than precise values of observations. However, the standard approaches that quantify uncertainty in machine learning cannot fully utilize censored labels. In this work, we adapt ensemble-based, Bayesian, and Gaussian models with tools to learn from censored labels by using the Tobit model from survival analysis. Our results demonstrate that despite the partial information available in censored labels, they are essential to accurately and reliably model the real pharmaceutical setting.

药物发现不确定性估计截断数据Tobit模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。