提出新算法DADO,利用科学设计中的可分解性提升离散对象生成效率。
Leveraging Discrete Function Decomposability for Scientific Design
- 基于变量间可分解结构,用图消息传递构建软因子化搜索分布
- 在蛋白设计等任务中显著加快收敛速度,优化效果优于传统方法
- 适合需要高效探索高维离散空间的科研设计场景
在人工智能驱动的科学与工程时代,我们常需根据用户指定属性在计算机中设计离散对象,例如设计能结合靶标的蛋白质、优化电路组件布局以降低延迟,或寻找特定性质的材料。给定属性预测模型后,通常通过在设计空间(如蛋白质序列空间)上训练生成模型,使设计集中在期望属性区域。分布优化——可形式化为估计分布算法或强化学习策略优化——旨在最大化目标函数的期望。然而,由于设计空间的组合性质,对离散设计分布进行优化通常极具挑战性。但许多科学应用中的属性预测器具有可分解性,即其可按设计变量分组因子化,理论上可提升优化效率。例如,蛋白质催化位点的氨基酸可能仅与其余部分弱相互作用即可实现最大催化活性。现有分布优化算法无法有效利用此类分解结构。本文提出并验证了一种新算法——分解感知分布优化(DADO),可利用设计变量上的任意连接树定义的可分解性,从而提升优化效率。核心在于采用软因子化的‘搜索分布’——一种学习得到的生成模型——结合图消息传递机制,在相关因子间协调优化,实现高效搜索空间导航。
原文摘要 · Abstract (English)
In the era of AI-driven science and engineering, we often want to design discrete objects in silico according to user-specified properties. For example, we may wish to design a protein to bind its target, arrange components within a circuit to minimize latency, or find materials with certain properties. Given a property predictive model, in silico design typically involves training a generative model over the design space (e.g., protein sequence space) to concentrate on designs with the desired properties. Distributional optimization$\unicode{x2013}$which can be formalized as an estimation of distribution algorithm or as reinforcement learning policy optimization$\unicode{x2013}$finds the generative model that maximizes an objective function in expectation. Optimizing a distribution over discrete-valued designs is in general challenging because of the combinatorial nature of the design space. However, many property predictors in scientific applications are decomposable in the sense that they can be factorized over design variables in a way that could in principle enable more effective optimization. For example, amino acids at a catalytic site of a protein may only loosely interact with amino acids of the rest of the protein to achieve maximal catalytic activity. Current distributional optimization algorithms are unable to make use of such decomposability structure. Herein, we propose and demonstrate use of a new distributional optimization algorithm, Decomposition-Aware Distributional Optimization (DADO), that can leverage any decomposability defined by a junction tree on the design variables, to make optimization more efficient. At its core, DADO employs a soft-factorized "search distribution"$\unicode{x2013}$a learned generative model$\unicode{x2013}$for efficient navigation of the search space, invoking graph message-passing to coordinate optimization across linked factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。