arXiv:2501.19153cs.LG2025-01被引 7

通过扩展测试时训练,提升化学分子生成的探索能力。

Test-Time Training Scaling Laws for Chemical Exploration in Drug Design

  • 采用多智能体测试时训练,模拟真实药物设计场景
  • 智能体数量按对数线性增长,显著提升分子多样性探索效率
  • 合作强化学习策略可进一步优化生成效果,适合药物研发人员

基于强化学习的化学语言模型在新分子设计中展现出潜力,但常因模式崩溃导致化学空间探索受限。受大语言模型中测试时训练(TTT)启发,我们提出将TTT扩展至化学语言模型,以增强分子空间探索能力。为此引入MolExp新基准,重点评估结构多样但生物活性相似的分子发现能力,模拟真实药物设计挑战。结果表明,通过增加独立强化学习智能体数量来扩展TTT,遵循对数线性缩放规律,显著提升探索效率(以MolExp衡量)。而延长TTT训练时间则带来边际收益递减,即使加入探索奖励也未能改善。我们还评估了协作强化学习策略,进一步提升探索效率。研究为生成式分子设计提供了可扩展框架,为优化AI驱动药物发现提供关键洞见。

原文摘要 · Abstract (English)

Chemical Language Models (CLMs) leveraging reinforcement learning (RL) have shown promise in de novo molecular design, yet often suffer from mode collapse, limiting their exploration capabilities. Inspired by Test-Time Training (TTT) in large language models, we propose scaling TTT for CLMs to enhance chemical space exploration. We introduce MolExp, a novel benchmark emphasizing the discovery of structurally diverse molecules with similar bioactivity, simulating real-world drug design challenges. Our results demonstrate that scaling TTT by increasing the number of independent RL agents follows a log-linear scaling law, significantly improving exploration efficiency as measured by MolExp. In contrast, increasing TTT training time yields diminishing returns, even with exploration bonuses. We further evaluate cooperative RL strategies to enhance exploration efficiency. These findings provide a scalable framework for generative molecular design, offering insights into optimizing AI-driven drug discovery.

分子生成强化学习药物设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。