用历史数据模拟实验顺序,大幅节省药物研发资源。
Case-Guided Sequential Assay Planning in Drug Discovery
- 基于相似案例构建隐式动态模型,实现贝叶斯更新。
- 真实任务中资源消耗降低92%,决策信心不下降。
- 适合数据多但无仿真器的药物研发场景。
在药物发现中,最优地规划实验序列是一个高风险、强不确定性且资源受限的决策问题。标准强化学习(RL)的主要障碍在于缺乏显式的环境模拟器或转移数据(s, a, s'),规划只能依赖静态的历史结果数据库。我们提出隐式贝叶斯马尔可夫决策过程(IBMDP),一种专为无模拟器环境设计的模型化强化学习框架。IBMDP通过使用相似的历史结果构建非参数信念分布,形成对转移动态的隐式建模。该机制支持随证据累积进行贝叶斯信念更新,并采用集成蒙特卡洛树搜索(ensemble MCTS)生成稳定策略,平衡信息获取与资源效率。我们在真实世界中枢神经系统(CNS)药物发现任务中验证了IBMDP,相比现有启发式方法,资源消耗最高减少92%,同时保持决策置信度。为进一步评估决策质量,我们在具有可计算最优策略的合成环境中进行基准测试,结果显示,相较于使用相同相似性模型的确定性值迭代方法,本框架与最优策略的对齐度显著更高,证明了其集成规划器的优势。IBMDP为数据丰富但缺乏模拟器的领域提供了实用的序列实验设计解决方案。
原文摘要 · Abstract (English)
Optimally sequencing experimental assays in drug discovery is a high-stakes planning problem under severe uncertainty and resource constraints. A primary obstacle for standard reinforcement learning (RL) is the absence of an explicit environment simulator or transition data $(s, a, s')$; planning must rely solely on a static database of historical outcomes. We introduce the Implicit Bayesian Markov Decision Process (IBMDP), a model-based RL framework designed for such simulator-free settings. IBMDP constructs a case-guided implicit model of transition dynamics by forming a nonparametric belief distribution using similar historical outcomes. This mechanism enables Bayesian belief updating as evidence accumulates and employs ensemble MCTS planning to generate stable policies that balance information gain toward desired outcomes with resource efficiency. We validate IBMDP through comprehensive experiments. On a real-world central nervous system (CNS) drug discovery task, IBMDP reduced resource consumption by up to 92\% compared to established heuristics while maintaining decision confidence. To rigorously assess decision quality, we also benchmarked IBMDP in a synthetic environment with a computable optimal policy. Our framework achieves significantly higher alignment with this optimal policy than a deterministic value iteration alternative that uses the same similarity-based model, demonstrating the superiority of our ensemble planner. IBMDP offers a practical solution for sequential experimental design in data-rich but simulator-poor domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。