用大模型引导进化搜索,让分子优化过程可解释且高效
MolEvolve: LLM-Guided Evolutionary Search for Interpretable Molecular Optimization
- 用大语言模型+蒙特卡洛树搜索,自动规划分子修改路径
- 在多个任务中超越基线,同时生成可读的化学推理链条
- 适合需要透明决策过程的药物研发人员
尽管深度学习在化学领域取得成功,但其应用受限于缺乏可解释性以及无法解决活性悬崖问题——微小结构变化引发性质剧烈波动。现有表征学习受相似性原则约束,难以捕捉这类结构-活性不连续性。为此,我们提出MolEvolve,一个将分子发现重构为自主前瞻规划问题的进化框架。不同于依赖人工特征或固定先验的传统方法,MolEvolve利用大语言模型(LLM)主动探索并演化一系列可执行的化学符号操作。通过使用LLM进行冷启动,结合蒙特卡洛树搜索(MCTS)与外部工具(如RDKit)在测试时进行规划,系统可自主发现最优演化路径。该过程生成透明的推理链,将复杂的结构变换转化为可理解的人类化学洞察。实验表明,MolEvolve的自主搜索不仅生成可读性强的化学见解,还在属性预测和分子优化任务中均优于基线方法。
原文摘要 · Abstract (English)
Despite deep learning's success in chemistry, its impact is hindered by a lack of interpretability and an inability to resolve activity cliffs, where minor structural nuances trigger drastic property shifts. Current representation learning, bound by the similarity principle, often fails to capture these structural-activity discontinuities. To address this, we introduce MolEvolve, an evolutionary framework that reformulates molecular discovery as an autonomous, look-ahead planning problem. Unlike traditional methods that depend on human-engineered features or rigid prior knowledge, MolEvolve leverages a Large Language Model (LLM) to actively explore and evolve a library of executable chemical symbolic operations. By utilizing the LLM to cold start and an Monte Carlo Tree Search (MCTS) engine for test-time planning with external tools (e.g. RDKit), the system self-discovers optimal trajectories autonomously. This process evolves transparent reasoning chains that translate complex structural transformations into actionable, human-readable chemical insights. Experimental results demonstrate that MolEvolve's autonomous search not only evolves transparent, human-readable chemical insights, but also outperforms baselines in both property prediction and molecule optimization tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。