arXiv:2609.04672cs.LG2026-09

无需预训练的分子指纹模型,性能全面领先基准方法。

WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding

论文配图:WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding
图 1 · 摘自论文原文
  • 用连续指纹将分子子结构分布到向量中,结合图距离旋转编码提升表示能力。
  • 在TDC ADMET 22项任务中10次夺冠,多项指标排名第一,且无需外部预训练。
  • 指纹可近乎无损重建分子结构,支持高效位置记忆,适合药物研发场景。

我们提出WEECFP,一种无参数的1024维连续分子指纹,将每个摩根子结构散列至约32个带符号位置;以及WEECFP-SuRGE,一种采用子结构旋转图距离编码(SuRGE)的Transformer架构,该编码基于分子最短路径图距离进行旋转。7模型融合的WEECFP-SuRGE Blend在TDC ADMET排行榜上平均回归排名最低,总分排名第2(仅次于预训练的MapLight+GNN),在无外部预训练方法中排名第一。在完整22项基准测试中,其在Pgp、Lipophilicity、CYP2D6 Substrate、Clearance Microsome和LD50上均获第一;其中不带SuRGE的融合版本在HIA任务也排名第一。在MoleculeNet上,该模型在ESOL、Lipophilicity和QM9三个回归任务上超越所有经典指纹基线。进一步显示,WEECFP分词几乎无损:贪婪重构建可恢复99.9%的分布内分子的规范SMILES(跨9个MoleculeNet数据集),在跨数据集验证(HIV→Lipophilicity)中恢复率达98.93%;三参考最远优先编码的图距离与真实距离相关系数为Pearson r = 0.901,实现O(S)级位置记忆精度。

原文摘要 · Abstract (English)

We introduce WEECFP, a parameter-free 1024-dimensional continuous molecular fingerprint that scatters each Morgan substructure across roughly thirty-two signed positions of a single vector, and WEECFP-SuRGE, a transformer architecture whose self-attention applies SuRGE (Substructure Rotary Graph-distance Encoding) -- a RoPE-like rotation parameterized by molecular shortest-path graph distance -- to WEECFP substructure tokens. A 7-model blend of this architecture (the WEECFP-SuRGE Blend) achieves the lowest average regression rank on the TDC ADMET leaderboard; is #2 overall on the TDC ADMET leaderboard (behind only pretrained MapLight+GNN), and is #1 overall among methods that use no external pretraining; takes leaderboard #1 finishes on Pgp, Lipophilicity, CYP2D6 Substrate, Clearance Microsome, and LD50 (with the WEECFP-NoSuRGE Blend separately reaching #1 on HIA) across the full 22-benchmark suite -- without any external pretraining. On MoleculeNet, WEECFP-SuRGE beats every classical-fingerprint baseline on 3 of 4 regression tasks (ESOL, Lipophilicity, QM9). We further show that WEECFP tokenization is near-lossless: a greedy overlap reconstruction recovers the exact canonical SMILES of 99.9% of in-distribution molecules across 9 MoleculeNet datasets and 98.93% of molecules in a cross-dataset holdout (HIV->Lipophilicity), and that a three-reference farthest-first encoding of graph distance correlates at Pearson r = 0.901 with the true pairwise distance, enabling O(S) positional memory at matching accuracy.

分子表示图神经网络指纹生成药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。