PrexSyn高效生成可合成分子,支持用户编程式定义性质目标。
Efficient and Programmable Exploration of Synthesizable Chemical Space
- 基于百亿级可合成路径数据训练的解码器模型,实现快速生成
- 能同时满足多条件性质查询,优化黑盒目标函数效率更高
- 适合药物研发与分子设计人员,尤其需高效筛选可合成分子场景
可合成化学空间的约束性给同时具备理想性质且可合成的分子采样带来重大挑战。本文提出PrexSyn,一种高效且可编程的可合成化学空间分子发现模型。PrexSyn基于解码器仅架构的Transformer,利用实时高通量C++数据生成引擎构建的百亿级可合成路径与分子性质数据流进行训练。大规模数据使PrexSyn在高速推理下近乎完美重建可合成化学空间,并学习性质与合成路径间的关联。基于已学映射,PrexSyn不仅能生成满足单一性质的分子,还可处理由逻辑运算符连接的复合性质查询,实现用户对生成目标的“编程”。此外,借助该性质查询能力,PrexSyn可通过迭代查询优化,以高于无合成限制基线的采样效率优化分子对抗黑盒目标函数,成为强大的通用分子优化工具。总体而言,PrexSyn在可合成化学空间覆盖、分子采样效率和推理速度上均达到新基准。
原文摘要 · Abstract (English)
The constrained nature of synthesizable chemical space poses a significant challenge for sampling molecules that are both synthetically accessible and possess desired properties. In this work, we present PrexSyn, an efficient and programmable model for molecular discovery within synthesizable chemical space. PrexSyn is based on a decoder-only transformer trained on a billion-scale datastream of synthesizable pathways paired with molecular properties, enabled by a real-time, high-throughput C++-based data generation engine. The large-scale training data allows PrexSyn to reconstruct the synthesizable chemical space nearly perfectly at a high inference speed and learn the association between properties and synthesizable molecules. Based on its learned property-pathway mappings, PrexSyn can generate synthesizable molecules that satisfy not only single-property conditions but also composite property queries joined by logical operators, thereby allowing users to ``program'' generation objectives. Moreover, by exploiting this property-based querying capability, PrexSyn can efficiently optimize molecules against black-box oracle functions via iterative query refinement, achieving higher sampling efficiency than even synthesis-agnostic baselines, making PrexSyn a powerful general-purpose molecular optimization tool. Overall, PrexSyn pushes the frontier of synthesizable molecular design by setting a new state of the art in synthesizable chemical space coverage, molecular sampling efficiency, and inference speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。