arXiv:2608.10480cs.AIcs.LG2026-08

让分子大模型读懂关键化学基团,提升性质预测准确性

Multi-Granular Rationale-Guided Molecular LLM for Property Prediction

论文配图:Multi-Granular Rationale-Guided Molecular LLM for Property Prediction
图 1 · 摘自论文原文
  • 用图神经网络识别影响性质的关键子结构,并生成带方向和排序的解释
  • 在8个MoleculeNet任务上超越通用模型,接近专用模型性能
  • 首次将GNN的归因结果作为证据输入大模型,适合药物研发人员使用

大语言模型广泛应用于化学任务,如分子性质预测,是药物发现的基础。现有分子大模型通过1D SMILES序列或2D分子图表示分子,但信息编码隐含,子结构贡献不透明。检索与增强方法虽可引入外部上下文,但无法反映化学家关注的内部子结构作用。本文提出MR-MoL,一种多粒度理由引导的分子大模型,直接提供由子结构驱动性质变化的证据。通过微调的图神经网络对子结构进行掩码评分,提取最具影响力的子结构,按重要性排序并标注方向,形成结构化理由序列,供大模型与SMILES及分子图一同读取。理由涵盖三个粒度层级:Murcko骨架及其侧链、BRICS片段、功能基团。这是首个将GNN衍生归因作为证据输入大模型的方法。在八个MoleculeNet任务中,MR-MoL在通用模型中表现最佳,并缩小了与专用模型的差距。五项诊断验证模型真正理解理由:其方向、排序与子结构均显著影响预测结果,且归因可复现已知结构-性质关系。

原文摘要 · Abstract (English)

Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph. Both encode molecular information implicitly, so the contribution of individual substructures remains opaque. Retrieval and augmentation methods add context, but from external sources. However, the cues chemists reason over are the internal substructures that drive a property up or down. We propose MR-MoL, a multi-granular rationale-guided molecular LLM that supplies this evidence directly. A fine-tuned GNN scores each substructure through masking, and the most influential ones are serialized as a ranked, direction-tagged rationale that the LLM reads alongside the SMILES sequence and molecular graph. The rationale spans three levels of granularity: Murcko scaffolds with their side chains, BRICS fragments, and functional groups. This is, to our knowledge, the first method to expose GNN-derived attributions to an LLM as evidence for property prediction. On eight MoleculeNet tasks, MR-MoL achieves the best overall results among generalist models and narrows the gap to specialist models tuned for each task. Five diagnostics further confirm that the model reads the rationale rather than merely benefiting from its presence. Its direction, rank, and substructure each shape the prediction, and its attributions reproduce known structure-property relationships.

分子生成大模型可解释性药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。