用大模型从多角度学习分子表示,提升性质预测效果。
$\text{M}^{2}$LLM: Multi-view Molecular Representation Learning with Large Language Models
- 从结构、任务、规则三视角动态融合分子表征
- 在多个分类与回归任务上达到当前最优性能
- 适合药物发现与材料科学中的分子智能研究
准确预测分子性质是化学、材料科学和药物发现中的关键挑战。传统分子表征方法(如指纹和图神经网络)通过提取分子结构特征取得先进成果,但常忽略数十年积累的语义与上下文知识。近期大型语言模型(LLMs)在科学领域展现出卓越的推理能力与先验知识,促使我们提出假设:若引导大模型从多视角推理,可生成丰富的分子表示。为此,我们提出 $ ext{M}^{2}$LLM,一个融合三种视角的多视图框架:分子结构视图、分子任务视图和分子规则视图。这些视图动态融合以适应不同任务需求。实验表明,$ ext{M}^{2}$LLM 在多个分类与回归基准上均实现领先性能。此外,我们验证了基于大模型的表征在两项核心能力下表现优异:一是利用编码能力生成分子嵌入,二是通过高级推理过程筛选与优化分子特征。
原文摘要 · Abstract (English)
Accurate molecular property prediction is a critical challenge with wide-ranging applications in chemistry, materials science, and drug discovery. Molecular representation methods, including fingerprints and graph neural networks (GNNs), achieve state-of-the-art results by effectively deriving features from molecular structures. However, these methods often overlook decades of accumulated semantic and contextual knowledge. Recent advancements in large language models (LLMs) demonstrate remarkable reasoning abilities and prior knowledge across scientific domains, leading us to hypothesize that LLMs can generate rich molecular representations when guided to reason in multiple perspectives. To address these gaps, we propose $\text{M}^{2}$LLM, a multi-view framework that integrates three perspectives: the molecular structure view, the molecular task view, and the molecular rules view. These views are fused dynamically to adapt to task requirements, and experiments demonstrate that $\text{M}^{2}$LLM achieves state-of-the-art performance on multiple benchmarks across classification and regression tasks. Moreover, we demonstrate that representation derived from LLM achieves exceptional performance by leveraging two core functionalities: the generation of molecular embeddings through their encoding capabilities and the curation of molecular features through advanced reasoning processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。