测试大模型对分子结构微小变化的适应能力,发现轻微改动就导致性能大幅下降。
Do LLMs Truly Generalize in the Molecular Domain? A Perturbation-Based Analysis

- 通过控制图编辑距离生成分子结构变体,测试模型泛化能力
- 单次结构修改即引发任务性能显著下降,信任区域极窄
- 上下文调优可部分缓解脆弱性,适合分子性质预测场景
大型语言模型在分子发现中展现出潜力,但其基于序列的随机性与化学空间严格的拓扑约束之间存在差距。这引发了关于分子大模型是否能超越序列表示所诱导的局部邻域泛化的疑问。为此,我们提出分子扰动框架,通过控制图编辑距离(GED)生成训练分子的语法有效结构变体,以探测分子大模型流形的规则性。分析表明,即使一次编辑也会导致常见分子任务性能显著下降,揭示了狭窄的局部信任区域和对结构变化的脆弱敏感性。由于相似分子通常具有相似性质,上下文调优(ICT)通过锚定在结构相似分子上的预测,成为缓解此类脆弱性的自然方式。实验还检验了ICT在可控结构扰动下的鲁棒性,结果表明其可部分扩展局部信任区域,为稳定分子大模型应对结构变异提供了有前景的方向。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently shown promise in molecular discovery, yet a gap remains between their probabilistic nature over discrete sequential tokens and the rigid topological constraints of chemical space. This raises the question of whether molecular LLMs can generalize beyond the local neighborhoods induced by their sequence-based representations. To systematically investigate this question, we introduce a Molecular Perturbation framework that generates syntax-valid structural variants of training molecules under controlled Graph Edit Distance (GED) to probe the manifold regularity of molecular LLMs. Our analysis shows that even a single edit can cause substantial performance drops on common molecular tasks, revealing a narrow local trust region and fragile sensitivity to structural changes. Since similar molecules tend to exhibit similar properties, In-Context Tuning (ICT), which anchors predictions on structurally similar molecules, offers a natural way to mitigate such fragility. Our experiments also examine whether ICT confers robustness under controlled structural perturbations, and the results suggest that it can partially expand the local trust region and offer a promising direction for stabilizing molecular LLMs against structural variation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。