通过分层结构-属性对齐,实现小数据下的分子生成与编辑。
Hierarchical Structure-Property Alignment for Data-Efficient Molecular Generation and Editing
- 分原子、片段和整体三个层级对分子结构与性质关系建模。
- 仅需少量数据即可训练,且在稀疏标注下仍能生成高质量分子。
- 适合药物发现领域需要精准调控分子性质的研究者使用。
基于属性约束的分子生成与编辑在人工智能驱动的药物发现中至关重要,但仍受制于两大挑战:(i) 分子结构与多属性间复杂关系难以捕捉;(ii) 分子属性覆盖范围窄且标注不全,削弱了基于属性模型的效果。为此,我们提出 HSPAG 框架,采用分层结构-属性对齐策略,将 SMILES 与分子属性视为互补模态,在原子、子结构和整体分子层面学习其关联。此外,通过骨架聚类选择代表性样本,并利用辅助变分自编码器(VAE)筛选困难样本,显著降低预训练所需数据量。同时引入属性相关性感知掩码机制与多样化扰动策略,提升稀疏标注下的生成质量。实验表明,HSPAG 能有效捕捉细粒度的结构-属性关系,并支持多属性约束下的可控生成。两个真实案例研究进一步验证了其编辑能力。
原文摘要 · Abstract (English)
Property-constrained molecular generation and editing are crucial in AI-driven drug discovery but remain hindered by two factors: (i) capturing the complex relationships between molecular structures and multiple properties remains challenging, and (ii) the narrow coverage and incomplete annotations of molecular properties weaken the effectiveness of property-based models. To tackle these limitations, we propose HSPAG, a data-efficient framework featuring hierarchical structure-property alignment. By treating SMILES and molecular properties as complementary modalities, the model learns their relationships at atom, substructure, and whole-molecule levels. Moreover, we select representative samples through scaffold clustering and hard samples via an auxiliary variational auto-encoder (VAE), substantially reducing the required pre-training data. In addition, we incorporate a property relevance-aware masking mechanism and diversified perturbation strategies to enhance generation quality under sparse annotations. Experiments demonstrate that HSPAG captures fine-grained structure-property relationships and supports controllable generation under multiple property constraints. Two real-world case studies further validate the editing capabilities of HSPAG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。