arXiv:2410.08829cs.LGcs.AI2024-10被引 1

通过挖掘分子语言模型的线性潜结构,提升可解释性与性能

Exploiting Latent Linearity in LLMs Improves Explainable Molecular Representation Learning

  • 将分子嵌入分解到与化学概念对齐的线性空间
  • 在下游任务中实现300倍加速、参数减少10万倍
  • 适合需要可解释性与高效推理的药物发现场景

大语言模型(LLMs)在药物发现和材料设计等分子领域展现出广泛适用性。分析其潜在表示对揭示内在机制、提升可解释性及推动下游性能至关重要。我们提出MoleX,一个简单有效的框架,将LLM中的分子嵌入分解至与化学概念对齐的潜空间,实现可解释的分子表示学习。进一步研究表明,这些高维嵌入可线性映射至化学一致的概念空间。分析表明,该线性结构与已知化学原理一致,揭示了科学应用中LLM表示的可解释性潜结构。应用于下游任务时,该线性结构同时提升预测与解释性能。大量实验表明,MoleX在准确性、可解释性和效率方面均优于现有方法:在大规模数据集上实现CPU推理速度提升300倍,参数量减少10万,显著优于原始LLMs。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated broad utility across molecular domains, spanning drug discovery and materials design. Analyzing LLMs' latent representations is crucial for elucidating their underlying mechanisms, improving explainability, and ultimately advancing downstream performance. We propose MoleX, a simple yet effective framework that decomposes molecular embeddings within LLM representations into a concept-aligned space for explainable molecular representation learning. We further show that these high-dimensional embeddings admit a linear mapping onto chemically consistent concepts. Our analysis suggests that the uncovered linearity aligns with established chemical principles, indicating a mechanistically explainable latent structure in LLM representations for scientific applications. When applied to downstream tasks, this latent linearity improves both predictive and explanatory performance. Extensive experiments demonstrate that MoleX outperforms existing approaches in accuracy, explainability, and efficiency, achieving CPU inference on large-scale datasets 300 times faster with 100,000 fewer parameters than LLMs.

可解释性分子表征语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。