揭示生成模型如何组织分子身份,发现其内部存在固定分区结构。
How Molecular Generative Models Organize Molecular Identity

- 通过反向追踪生成过程,显式分析分子身份的内部布局。
- 模型生成区域呈分块常数分布,边界具有粗到细的重复特征。
- 组织方式受表示、解码随机性等因素影响,需谨慎评估化学可导航性。
物质生成模型通常以输出表示的采样器来评估,其潜在空间常被用作化学空间的导航代理。然而,这些模型如何在内部排列离散化学身份仍不清楚。本文通过显式化分子身份并将其回溯至生成过程,探查生成同一对象的区域,揭示了训练模型的内部谱系:一个决定模型能生成哪些对象(新或旧)的固定分区。在三种分子生成架构中,该谱系呈现由反复出现的粗到细边界分隔的分段常数区域。其组织方式依赖于所探查的表示、身份约定、解码随机性和坐标比较度量。训练过程中,局部化学组织趋于稳定,但每个邻域内代表的分子身份数量仍在持续变化。因此,在将生成空间视为化学可导航之前,必须对其内部组织进行表征,而非假设。
原文摘要 · Abstract (English)
Generative models for matter are often evaluated as samplers over output representations, and their latent spaces are commonly used as proxies for navigating chemical space. Much less is known about how these models internally arrange discrete chemical identities within those representations. We study this arrangement by making molecular identity explicit and pulling it back through the generative process. Through these pullbacks we probe the regions that generate the same object, exposing the trained model's internal repertoire: a fixed partition that determines which objects (novel or not) the model can produce. Across three molecular generative architectures, we find that this repertoire is arranged into piecewise-constant regions separated by recurring coarse-to-fine boundaries. Its organization depends on the representation probed, the identity convention, decoder stochasticity, and the metric used to compare coordinates. During training, local chemical organization stabilizes while the number of distinct molecular identities represented within each neighborhood continues to change. Internal organization must therefore be characterized, rather than assumed, before a generative space can be treated as chemically navigable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。