arXiv:2507.17876cs.LGphysics.chem-ph2025-07被引 1

用负样本训练模型生成正向分子,无需正样本数据。

Look the Other Way: Designing 'Positive' Molecules with Negative Data via Task Arithmetic

  • 通过负样本学习属性方向,反向移动生成理想分子。
  • 在33组实验中生成更多样且成功的分子设计。
  • 适合小分子、蛋白质等多场景的零样本与少样本设计。

理想分子(即‘正向’分子)数据稀缺是生成式分子设计的核心瓶颈。为突破此限制,本文提出分子任务算术:仅使用丰富且多样的负样本训练模型,学习‘属性方向’,无需正样本标签即可反向移动生成正向分子。在33组涵盖小分子、蛋白质、不同模型架构与规模的设计实验中,该方法生成的分子在多样性和成功率上均优于基于正样本训练的模型。此外,在双目标与少样本设计任务中,分子任务算术持续提升设计多样性,同时保持良好的蛋白对接评分等复杂性质。其简单性、数据高效性与优异性能使其有望成为从头分子设计的主流迁移学习策略。

原文摘要 · Abstract (English)

The scarcity of molecules with desirable properties (i.e., `positive' molecules) is an inherent bottleneck for generative molecule design. To sidestep such obstacle, here we propose molecular task arithmetic: training a model on diverse and abundant negative examples to learn 'property directions' - without accessing any positively labeled data - and moving models in the opposite property directions to generate positive molecules. When analyzed on 33 design experiments with distinct molecular entities (small molecules, proteins), model architectures, and scales, molecular task arithmetic generated more diverse and successful designs than models trained on positive molecules in general. Moreover, we employed molecular task arithmetic in dual-objective and few-shot design tasks. We find that molecular task arithmetic can consistently increase the diversity of designs while maintaining desirable complex design properties, such as good docking scores to a protein. With its simplicity, data efficiency, and performance, molecular task arithmetic bears the potential to become the de facto transfer learning strategy for de novo molecule design.

分子生成负样本迁移学习零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。