arXiv:2604.23061cs.LGcs.AI2026-04

用强化学习让大模型更精准地优化药物分子,兼顾多个目标且不破坏骨架结构。

C-MORAL: Controllable Multi-Objective Molecular Optimization with Reinforcement Alignment for LLMs

论文配图:C-MORAL: Controllable Multi-Objective Molecular Optimization with Reinforcement Alignment for LLMs
图 1 · 摘自论文原文
  • 通过分组相对优化与非线性奖励聚合,稳定处理多目标冲突的分子设计。
  • 在域内任务中成功率48.9%,域外任务39.5%,同时保持骨架相似性。
  • 适合需要多目标协同优化的药物分子生成研究者使用。

大语言模型在分子优化中展现潜力,但难以对齐具有选择性和竞争性的药物设计约束。我们提出C-MORAL,一种基于强化学习的后训练框架,实现可控的多目标分子优化。该方法结合分组相对优化、异构目标属性得分对齐以及瓶颈敏感的非线性奖励聚合,提升在竞争性分子性质下的稳定性。在C-MuMOInstruct和S²-Bench MolOpt数据集上的实验表明,C-MORAL在两个基准上均优于现有方法。在C-MuMOInstruct上,域内任务成功优化率(SOR)达48.9%,域外任务为39.5%,同时保持骨架相似性;在S²-Bench MolOpt上,对LogP、MR和QED优化任务均取得最强效果。结果表明,C-MORAL是有效对齐分子LLM与连续、受限设计目标的方法。代码与模型已开源:https://github.com/Rwigie/C-MORAL。

原文摘要 · Abstract (English)

Large language models (LLMs) show promise for molecular optimization, but aligning them with selective and competing drug-design constraints remains challenging. We propose C-Moral, a reinforcement learning post-training framework for controllable multi-objective molecular optimization. C-Moral combines group-based relative optimization, property score alignment for heterogeneous objectives, and bottleneck-sensitive non-linear reward aggregation to improve stability across competing molecular properties. Experiments on C-MuMOInstruct and S$^2$-Bench MolOpt show that C-Moral achieves the best performance among compared methods on both benchmarks. On C-MuMOInstruct, C-Moral achieves the best Success Optimized Rate (SOR) of 48.9\% on in-domain tasks and 39.5\% on out-of-domain tasks while preserving scaffold similarity. On S$^2$-Bench MolOpt, it also achieves the strongest results across LogP, MR, and QED optimization tasks. These results suggest that C-Moral is an effective way to align molecular LLMs with continuous and constrained molecular design objectives. Our code and models are publicly available at https://github.com/Rwigie/C-MORAL.

分子生成强化学习多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。