用可学习的碎片化方法,让分子表示更精准。
FragmentNet: Adaptive Graph Fragmentation for Graph-to-Sequence Molecular Representation Learning
- 自适应生成化学有效碎片,替代固定规则分词
- 在多个属性预测任务中,碎片级预训练提升性能
- 适合需要精细分子结构建模的研究者
分子表示学习通常将分子拆分为单个原子,或使用固定规则的片段分解方式,限制了对有意义化学子结构上下文的捕捉能力。我们提出FragmentNet,一种基于新型自适应、可学习分词器的图到序列模型,该分词器可将分子图分解为可调节粒度的化学有效片段,并引入化学感知的空间位置编码以保留分子拓扑结构。借鉴自然语言处理中的掩码预训练策略,我们在化学上有意义的片段层面进行掩码与重建,而非原子级别。在多个属性预测基准上评估发现,以片段粒度进行预训练能显著提升多数任务的下游表现,证明了分词粒度是分子表示学习中的关键设计因素。
原文摘要 · Abstract (English)
Molecular representation learning methods typically tokenize molecules as individual atoms or use rigid, rule-based fragment decompositions, limiting their ability to capture meaningful chemical substructure context. We introduce FragmentNet, a graph-to-sequence model built around a novel adaptive, learned tokenizer that decomposes molecular graphs into chemically valid fragments of adjustable granularity, complemented by chemically aware spatial positional encodings that preserve molecular topology in the resulting sequence. Extending masked pre-training strategies from natural language processing to the molecular domain, we mask and reconstruct molecules at the level of chemically meaningful fragments rather than individual atoms. Evaluating across multiple property prediction benchmarks, we find that pre-training at fragment granularity leads to improved downstream performance on the majority of tasks, demonstrating that tokenization granularity is an important design choice for molecular representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。