用稀疏自编码器揭示化学语言模型如何逐层构建分子表示
What Does a Chemical Language Model Know About Molecules?

- 通过稀疏自编码器分析模型各层的潜在表征机制
- 早期层依赖位置追踪,后期层编码子结构与药理特征
- 非标准SMILES比无效SMILES更破坏表示一致性,适合可解释性研究
化学语言模型(cLMs)通常被认为仅学习表面语法模式,而非深层分子语义。本文采用稀疏自编码器(SAEs)对仅编码器结构的cLM MolFormer进行机制分析,揭示其在不同层级上如何构建分子表征。结果发现:早期层依赖位置追踪潜在变量解析分子语法;后期层则编码原子在子结构中的角色及药理相关特征。此外,我们发现非标准SMILES比无效SMILES引发更剧烈的表征偏移,根源在于位置潜变量的扰动在层间传播。为支持后续探索,我们开发了InterMol,一个交互式可视化工具,用于观察分子字符串与结构上的SAE激活情况。
原文摘要 · Abstract (English)
Chemical language models (cLMs) are widely assumed to learn surface-level syntactic patterns rather than learning meaningful molecular semantics. Here, we apply sparse autoencoders (SAEs) to MolFormer, an encoder-only cLM, to mechanistically examine how molecular representations are built across layers. We discover that early layers rely on position-tracking latents to parse molecular grammar, while later layers encode atom-in-substructure and pharmacologically relevant features. Additionally, we show that non-canonical SMILES produce more disruptive representation shifts than invalid SMILES, driven by position-latent disruption propagating across layers. To support further exploration, we develop InterMol, an interactive visualizer for SAE activations on molecular strings and structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。