从大规模分子数据中高效选出能覆盖多目标空间的代表性样本,加速药物研发。
Multi-Objective Coverage via Constraint Active Search
- 基于高斯过程和置信上界,主动搜索能代表多目标空间的分子样本。
- 在新冠与癌症蛋白靶点数据集上,5个目标下均优于现有方法。
- 适合需要快速筛选化学多样性且满足安全约束的药物设计场景。
本文提出多目标覆盖(MOC)新问题:从可行空间中找出少量代表性样本,使其预测结果能广泛覆盖多目标区域。该问题对药物发现与材料设计等实际应用至关重要,因代表样本可显著加快科学探索速度。现有方法无法直接应用,因其要么关注样本空间覆盖,要么聚焦于帕累托前沿优化。而化学多样性样本常产生相同目标表现,且安全约束通常作用于目标值本身。为此,我们提出MOC-CAS算法,采用基于置信上界的采集函数,依据高斯过程后验预测选择乐观样本。为实现高效优化,我们设计了硬可行性测试的平滑松弛,并推导出近似优化器。在大规模蛋白-靶点数据集(分别用于SARS-CoV-2与癌症)上,每个数据集评估五个由SMILES特征衍生的目标,实验表明MOC-CAS在多个指标上显著优于基线方法。
原文摘要 · Abstract (English)
In this paper, we formulate the new multi-objective coverage (MOC) problem where our goal is to identify a small set of representative samples whose predicted outcomes broadly cover the feasible multi-objective space. This problem is of great importance in many critical real-world applications, e.g., drug discovery and materials design, as this representative set can be evaluated much faster than the whole feasible set, thus significantly accelerating the scientific discovery process. Existing works cannot be directly applied as they either focus on sample space coverage or multi-objective optimization that targets the Pareto front. However, chemically diverse samples often yield identical objective profiles, and safety constraints are usually defined on the objectives. To solve this MOC problem, we propose a novel search algorithm, MOC-CAS, which employs an upper confidence bound-based acquisition function to select optimistic samples guided by Gaussian process posterior predictions. For enabling efficient optimization, we develop a smoothed relaxation of the hard feasibility test and derive an approximate optimizer. Compared to the competitive baselines, we show that our MOC-CAS empirically achieves superior performances across large-scale protein-target datasets for SARS-CoV-2 and cancer, each assessed on five objectives derived from SMILES-based features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。