用稀疏自编码器挖掘化学大模型的隐藏知识,揭示其内部表征机制。
Unveiling Latent Knowledge in Chemistry Language Models through Sparse Autoencoders
- 通过稀疏自编码器解析化学大模型的潜在特征,实现可解释性分析。
- 发现潜空间特征与分子结构、物化性质和药理类别存在明确关联。
- 方法通用性强,适合研究化学AI系统内部机制,助力药物材料研发。
自机器学习兴起以来,可解释性始终是核心挑战,尤其在药物与材料发现等高风险应用中愈发紧迫。近年来,大型语言模型架构的进步催生了具备强大分子属性预测与生成能力的化学语言模型(CLMs),但其内部如何表征化学知识仍不清晰。本文将稀疏自编码器技术扩展至化学语言模型,应用于面向材料的基础模型FM4M SMI-TED,提取出具有语义意义的潜在特征,并分析其在多种分子数据集上的激活模式。结果表明,模型编码了丰富的化学概念体系,特定潜空间特征与分子结构片段、物化性质及药理类别存在显著对应关系。该方法为揭示化学导向人工智能系统的潜在知识提供了通用框架,对基础理解与实际应用均有重要意义,有望加速计算化学研究进程。
原文摘要 · Abstract (English)
Since the advent of machine learning, interpretability has remained a persistent challenge, becoming increasingly urgent as generative models support high-stakes applications in drug and material discovery. Recent advances in large language model (LLM) architectures have yielded chemistry language models (CLMs) with impressive capabilities in molecular property prediction and molecular generation. However, how these models internally represent chemical knowledge remains poorly understood. In this work, we extend sparse autoencoder techniques to uncover and examine interpretable features within CLMs. Applying our methodology to the Foundation Models for Materials (FM4M) SMI-TED chemistry foundation model, we extract semantically meaningful latent features and analyse their activation patterns across diverse molecular datasets. Our findings reveal that these models encode a rich landscape of chemical concepts. We identify correlations between specific latent features and distinct domains of chemical knowledge, including structural motifs, physicochemical properties, and pharmacological drug classes. Our approach provides a generalisable framework for uncovering latent knowledge in chemistry-focused AI systems. This work has implications for both foundational understanding and practical deployment; with the potential to accelerate computational chemistry research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。