揭示化学模型如何从分子字符串中学会手性语义,突破表层模式依赖。
From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models

- 通过自回归Transformer模型追踪训练过程,发现手性信息学习存在突然跃升
- 手性识别准确率在长期停滞后骤升,表明学习复杂性非单纯模型容量所致
- 编码器主导手性表示重构,关键注意力头可解释为手性敏感机制
理解化学语言模型(CLMs)如何从分子字符串表征中学习化学语义,而非仅依赖表面字符串模式,是化学表示学习与机器学习在化学领域应用中的核心问题。手性是一个严峻的测试案例:对映体在药理活性和毒性上可能差异巨大,但现有模型常难以可靠区分手性构型。本文提出Pan-CORE(泛化学多尺度表征引擎),一类基于自回归Transformer的编码器-解码器模型,用于SMILES翻译,并通过高时间分辨率检查点分析,研究手性信息在训练中的学习过程。所有测试的Pan-CORE变体均显示,手性标记准确率在长期平台期后出现明显跃升,表明手性学习停滞不能仅由模型容量解释,而是反映了手性约束的复杂性。注意力动态、残差流轨迹和潜在空间几何分析支持编码器中心机制:手性标记表示经历短暂失稳与重建,表现为向量范数和方向稳定性呈V形下降与恢复,并伴随潜在空间中手性分子表示的显著重组。编码器-解码器交叉评估进一步支持该转变的编码器中心特性,针对性注意力头消融实验识别出少数手性敏感头,其移除即使在完全训练模型中仍会显著降低手性标记准确率。这些发现表明,SMILES翻译可作为解析CLMs中语义涌现的实用实验系统,对可解释的化学表示学习具有启示意义。
原文摘要 · Abstract (English)
Understanding how chemical language models (CLMs) learn chemical meaning from molecular string representations, rather than only surface-level string patterns, is an important question in chemical representation learning and machine learning for chemistry. Chirality provides a demanding test case: enantiomers can differ greatly in pharmacological activity and toxicity, yet CLMs often struggle to distinguish chiral configurations reliably. Here we present Pan-CORE (Pan-Chemical Omniscale Representation Engine), a family of autoregressive Transformer-based encoder-decoder models for SMILES translation, and use high-temporal-resolution checkpoint analysis to investigate how chiral information is learned during training. Across all tested Pan-CORE variants, we observe a reproducible jump-up in which chiral-token accuracy rises abruptly after a long plateau, suggesting that chiral learning stagnation is not explained by model capacity alone and instead reflects the complexity of chiral constraints. Analyses of attention dynamics, residual-stream trajectories, and latent-space geometry support an encoder-centered mechanism in which chiral-token representations undergo transient destabilization and reconstruction, seen as a V-shaped drop and recovery in vector norm and directional stability, together with a clear reorganization of chiral molecular representations in the latent space. Encoder-decoder cross-evaluation further supports the encoder-centered nature of the transition, and targeted attention-head ablation identifies a small set of chiral-sensitive heads whose removal selectively reduces chiral-token accuracy even in the fully trained model. These findings show that SMILES translation can serve as a useful experimental system for mechanistic analysis of semantic emergence in CLMs, with implications for interpretable chemical representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。