构建首个支持推理模型从头生成分子的评测基准
MolRGen: A Training and Evaluation Setting for De Novo Molecular Generation with Reasonning Models
- 设计可实时计算奖励的分子生成环境,支持无参照物的全新分子设计
- 涵盖5万条多目标优化提示,包含对接评分与理化性质等指标
- 适用于研究验证器驱动的推理与强化学习在药物设计中的应用
近期基于推理的大语言模型在可验证任务中表现优异,但在从头分子生成方面仍受限于缺乏无需参考分子即可计算奖励的训练环境。本文提出MolRGen,一个用于训练和评估推理型大模型进行从头分子生成的基准与分子验证器。MolRGen包含约4,500个蛋白口袋靶点,生成50,000个结合评分与分子性质(如QED、合成可及性、logP及理化描述符)相结合的多目标优化提示。不同于基于图像描述或分子编辑的基准,MolRGen在生成时即时计算奖励,评估完全从零提出的分子。我们测试了通用与化学专用的开源大模型,并引入多样性感知的top-k度量,以衡量模型生成多样高分分子的能力。最后,利用验证器对128B参数模型进行GRPO微调,虽提升性能但带来多样性与探索性的权衡。MolRGen为研究基于验证器的推理与强化学习在分子设计中的应用提供了可扩展的实验平台。
原文摘要 · Abstract (English)
Recent reasoning-based large language models have shown strong performance on tasks with verifiable outcomes, but their use in de novo molecular generation remains limited by the lack of training environments where rewards can be computed without reference molecules. We introduce MolRGen, a benchmark and molecular verifier for training and evaluating reasoning LLMs on de novo molecular generation. MolRGen contains approximately 4,500 protein-pocket targets, resulting in 50k multi-objective optimization prompts combining docking scores with molecular properties such as QED, synthetic accessibility, logP, and physicochemical descriptors. Unlike caption-based generation or molecule-editing benchmarks, MolRGen evaluates molecules proposed from scratch by computing rewards at generation time. We benchmark general-purpose and chemistry-specialized open-source LLMs and introduce a diversity-aware top-k metric to measure whether models can generate a diverse set of high-scoring molecules. Finally, we use the verifier to fine-tune a 128B LLM with GRPO, showing improved performance, at the cost of a diversity-exploitation trade-off. MolRGen provides a scalable testbed for studying verifier-based reasoning and reinforcement learning in molecular design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。