让分子生成模型像调音一样快速切换属性组合,无需重新训练。
Mol-MoE: Training Preference-Guided Routers for Molecule Generation
- 用偏好学习训练路由器,动态选择专家模型实现属性调节
- 在多个化学性质上同时优化,生成分子质量优于现有方法
- 适合药物设计中快速探索不同属性权衡的场景
近期语言模型的发展使分子生成可视为序列建模问题。然而,现有方法多依赖单目标强化学习,难以应对真实药物设计中多个相互竞争属性需协同优化的需求。传统多目标强化学习(MORL)方法需为每种新目标组合重新训练,难以实现快速权衡探索。为此,我们提出 Mol-MoE,一种混合专家(MoE)架构,可在不重训的情况下实现测试时高效调控分子生成。其核心是基于偏好的路由器训练目标,促使路由器根据用户指定的权衡方式组合专家。该方法提升了测试时对化学性质空间的探索灵活性,支持快速权衡分析。在与先进方法的对比中,Mol-MoE 在样本质量和可调控性方面均表现更优。
原文摘要 · Abstract (English)
Recent advances in language models have enabled framing molecule generation as sequence modeling. However, existing approaches often rely on single-objective reinforcement learning, limiting their applicability to real-world drug design, where multiple competing properties must be optimized. Traditional multi-objective reinforcement learning (MORL) methods require costly retraining for each new objective combination, making rapid exploration of trade-offs impractical. To overcome these limitations, we introduce Mol-MoE, a mixture-of-experts (MoE) architecture that enables efficient test-time steering of molecule generation without retraining. Central to our approach is a preference-based router training objective that incentivizes the router to combine experts in a way that aligns with user-specified trade-offs. This provides improved flexibility in exploring the chemical property space at test time, facilitating rapid trade-off exploration. Benchmarking against state-of-the-art methods, we show that Mol-MoE achieves superior sample quality and steerability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。