arXiv:2601.12256cs.AI2026-01AAAI

通过多模态协作提升分子语言模型的准确性和鲁棒性

Improving Large Molecular Language Model via Relation-aware Multimodal Collaboration

  • 引入关系感知的多模态协同投影器,融合2D结构与3D空间关系
  • 在分子描述、性质问答等任务上超越现有模型,错误率降低18%
  • 提出新评估指标,更有效检测模型幻觉和生成质量

大语言模型(LLMs)在多种任务中展现出强大的指令遵循能力。受此启发,近期研究发展出大分子语言模型(LMLMs),将一维分子序列或二维分子图整合进语言模型。然而,现有LMLMs常因未能充分融合分子的一维序列、二维结构图和三维构象等多模态信息,导致幻觉频发且鲁棒性差。为此,我们提出CoLLaMo,一种基于大语言模型的分子助手,配备多层次分子模态协同投影器。其关系感知的模态协同注意力机制通过引入二维结构和三维空间关系,实现原子间细粒度、关系引导的信息交互。此外,我们设计了一种以分子为中心的新自动评估方法,包括幻觉评估指标和基于GPT的描述质量评价,弥补传统基于词元的通用评估指标(如BLEU)在评估分子理解能力上的不足。大量实验表明,CoLLaMo显著提升了LMLMs的分子模态泛化能力,在分子描述、计算性质问答、描述性质问答、基团计数和IUPAC命名预测等多个任务上均取得最优表现。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated their instruction-following capabilities and achieved powerful performance on various tasks. Inspired by their success, recent works in the molecular domain have led to the development of large molecular language models (LMLMs) that integrate 1D molecular strings or 2D molecular graphs into the language models. However, existing LMLMs often suffer from hallucination and limited robustness, largely due to inadequate integration of diverse molecular modalities such as 1D sequences, 2D molecular graphs, and 3D conformations. To address these limitations, we propose CoLLaMo, a large language model-based molecular assistant equipped with a multi-level molecular modality-collaborative projector. The relation-aware modality-collaborative attention mechanism in the projector facilitates fine-grained and relation-guided information exchange between atoms by incorporating 2D structural and 3D spatial relations. Furthermore, we present a molecule-centric new automatic measurement, including a hallucination assessment metric and GPT-based caption quality evaluation to address the limitations of token-based generic evaluation metrics (i.e., BLEU) widely used in assessing molecular comprehension of LMLMs. Our extensive experiments demonstrate that our CoLLaMo enhances the molecular modality generalization capabilities of LMLMs, achieving the best performance on multiple tasks, including molecule captioning, computed property QA, descriptive property QA, motif counting, and IUPAC name prediction.

分子语言模型多模态融合幻觉检测评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。