用拓扑瓶颈压缩蛋白结构,高效生成药物分子。
SiDGen: Structure-informed Diffusion for Generative modeling of Ligands for Proteins
- 通过软分配机制将残基信息压缩为低维瓶颈,降低计算开销。
- 在CrossDocked2020和DUD-E上达当前最佳性能,内存消耗显著下降。
- 适合需要高通量、结构感知药物生成的研究者使用。
基于结构的药物设计面临可扩展性与保真度的矛盾:充分考虑口袋几何信息虽准确但计算成本高,常随蛋白长度呈二次方(O(L²))或更坏增长;而仅依赖序列的条件生成则可能丢失关键相互作用结构。我们提出SiDGen,一种结构感知的离散扩散生成框架,通过拓扑信息瓶颈(TIB)解决这一权衡。SiDGen利用学习得到的软分配机制,将残基级蛋白表征压缩为紧凑瓶颈,在粗网格上进行下游成对计算(O(L²/s²)),大幅降低内存与计算开销,同时保持生成精度。该方法在CrossDocked2020和DUD-E基准测试中达到当前最优表现,并显著减少成对张量内存占用。SiDGen弥合了序列效率与口袋感知之间的差距,为高通量结构驱动药物发现提供可扩展路径。
原文摘要 · Abstract (English)
Structure-based drug design (SBDD) faces a fundamental scaling fidelity dilemma: rich pocket-aware conditioning captures interaction geometry but can be costly, often scales quadratically ($O(L^2)$) or worse with protein length ($L$), while efficient sequence-only conditioning can miss key interaction structure. We propose SiDGen, a structure-informed discrete diffusion framework that resolves this trade-off through a Topological Information Bottleneck (TIB). SiDGen leverages a learned, soft assignment mechanism to compress residue-level protein representations into a compact bottleneck enabling downstream pairwise computations on the coarse grid ($O(L^2/s^2)$). This design reduces memory and computational cost without compromising generative accuracy. Our approach achieves state-of-the-art performance on CrossDocked2020 and DUD-E benchmarks while significantly reducing pairwise-tensor memory. SiDGen bridges the gap between sequence-based efficiency and pocket-aware conditioning, offering a scalable path for high-throughput structure-based discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。