arXiv:2409.19740cs.LGq-bio.QM2024-09

用字节级分词+强化学习提升分子生成效率与质量

When Molecular GAN Meets Byte-Pair Encoding

  • 用字节对编码替代字符分词,更好捕捉复杂分子结构
  • 结合强化学习,生成分子的有效性、新颖性、多样性显著提升
  • 适合药物发现领域研究者快速生成高质量候选分子

深度生成模型如生成对抗网络(GAN)在从头生成新药分子方面至关重要。然而,传统基于字符的分词器常难以识别分子数据中的新颖复杂子结构。相比之下,其他分词方法表现更优。本研究提出一种集成字节级字节对编码分词器的分子GAN,并采用强化学习提升从头分子生成效果。生成器作为执行者生成SMILES字符串,判别器作为评价者评估其质量。该模型还引入创新奖励机制以提高计算效率。实验通过有效性、唯一性、新颖性和多样性评估,结合可视化分析,充分验证了该GAN的有效性。

原文摘要 · Abstract (English)

Deep generative models, such as generative adversarial networks (GANs), are pivotal in discovering novel drug-like candidates via de novo molecular generation. However, traditional character-wise tokenizers often struggle with identifying novel and complex sub-structures in molecular data. In contrast, alternative tokenization methods have demonstrated superior performance. This study introduces a molecular GAN that integrates a byte level byte-pair encoding tokenizer and employs reinforcement learning to enhance de novo molecular generation. Specifically, the generator functions as an actor, producing SMILES strings, while the discriminator acts as a critic, evaluating their quality. Our molecular GAN also integrates innovative reward mechanisms aimed at improving computational efficiency. Experimental results assessing validity, uniqueness, novelty, and diversity, complemented by detailed visualization analysis, robustly demonstrate the effectiveness of our GAN.

分子生成GAN强化学习SMILES

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。