arXiv:2512.14418cs.LG2025-12

构建全覆盖分子空间数据集,实现模型收敛学习与强泛化能力。

Dual-Axis RCCL: Representation-Complete Convergent Learning for Organic Chemical Space

  • 设计双轴表示法,融合局部价态与环/笼拓扑编码
  • 基于FD25数据集训练模型,预测误差仅1.0 kcal/mol MAE
  • 适合分子建模、材料发现等领域追求高泛化的研究者

机器学习正深刻改变分子与材料建模;然而面对化学空间(10^30-10^60)的巨大规模,模型能否在该空间实现收敛学习仍是开放问题。本文提出双轴表示完备的收敛学习(RCCL)策略,结合基于价键理论的图卷积网络(GCN)对局部价态环境的编码,以及无桥图(NBG)对环/笼拓扑的编码,提供化学空间覆盖率的量化度量。该框架形式化了表示完备性,为构建支持大模型收敛学习的数据集奠定原则基础。基于此,我们构建了FD25数据集,系统覆盖13,302种局部价态单元和165,726种环/笼拓扑,实现含H/C/N/O/F元素有机分子的近似完全组合覆盖。基于FD25训练的图神经网络展现出表示完备的收敛学习与强分布外泛化能力,在外部基准上整体预测误差约为1.0 kcal/mol MAE。结果建立了分子表示、结构完备性与模型泛化之间的定量联系,为可解释、可迁移、数据高效的分子智能提供基础。

原文摘要 · Abstract (English)

Machine learning is profoundly reshaping molecular and materials modeling; however, given the vast scale of chemical space (10^30-10^60), it remains an open scientific question whether models can achieve convergent learning across this space. We introduce a Dual-Axis Representation-Complete Convergent Learning (RCCL) strategy, enabled by a molecular representation that integrates graph convolutional network (GCN) encoding of local valence environments, grounded in modern valence bond theory, together with no-bridge graph (NBG) encoding of ring/cage topologies, providing a quantitative measure of chemical-space coverage. This framework formalizes representation completeness, establishing a principled basis for constructing datasets that support convergent learning for large models. Guided by this RCCL framework, we develop the FD25 dataset, systematically covering 13,302 local valence units and 165,726 ring/cage topologies, achieving near-complete combinatorial coverage of organic molecules with H/C/N/O/F elements. Graph neural networks trained on FD25 exhibit representation-complete convergent learning and strong out-of-distribution generalization, with an overall prediction error of approximately 1.0 kcal/mol MAE across external benchmarks. Our results establish a quantitative link between molecular representation, structural completeness, and model generalization, providing a foundation for interpretable, transferable, and data-efficient molecular intelligence.

分子建模图神经网络数据集构建泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。