arXiv:2505.10726cs.LGcs.AI2025-05NeurIPS

提出GRIN方法,让聚合物表示不随重复单元数量变化

Learning Repetition-Invariant Representations for Polymer Informatics

  • 用最大生成树对齐和重复单元增强,实现结构一致性
  • 三重复单元是获得最优不变表示的最小增强量
  • 在均聚物与共聚物上表现优于现有方法,泛化性强

聚合物是由重复结构单元(单体)构成的大分子,在能源存储、建筑、医疗和航空航天等领域广泛应用。然而,现有图神经网络方法虽对小分子有效,仅建模单个重复单元,无法为具有不同单元数量的真实聚合物结构生成一致的向量表示。为此,我们提出图重复不变性(GRIN)方法,学习对重复单元数量不变的聚合物表示。GRIN结合基于图的最大生成树对齐与重复单元增强,确保结构一致性。我们从模型和数据角度提供重复不变性的理论保证,证明三个重复单元是最小所需增强量以实现最优不变表示学习。GRIN在均聚物与共聚物基准测试中超越当前最优基线,学习到稳定且重复不变的表示,能有效泛化至未见长度的聚合物链。

原文摘要 · Abstract (English)

Polymers are large macromolecules composed of repeating structural units known as monomers and are widely applied in fields such as energy storage, construction, medicine, and aerospace. However, existing graph neural network methods, though effective for small molecules, only model the single unit of polymers and fail to produce consistent vector representations for the true polymer structure with varying numbers of units. To address this challenge, we introduce Graph Repetition Invariance (GRIN), a novel method to learn polymer representations that are invariant to the number of repeating units in their graph representations. GRIN integrates a graph-based maximum spanning tree alignment with repeat-unit augmentation to ensure structural consistency. We provide theoretical guarantees for repetition-invariance from both model and data perspectives, demonstrating that three repeating units are the minimal augmentation required for optimal invariant representation learning. GRIN outperforms state-of-the-art baselines on both homopolymer and copolymer benchmarks, learning stable, repetition-invariant representations that generalize effectively to polymer chains of unseen sizes.

聚合物图神经网络不变表示机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。