arXiv:2508.06917q-bio.QMcs.AI2025-08中稿 · ACMMM 2025被引 1

用跨视图前缀融合分子拓扑与空间结构,提升大模型理解能力

CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs

  • 通过引导重采样将分子图和图像视图转为固定长度前缀
  • 在分子描述生成、命名预测等任务上显著优于基线方法
  • 适合需要精准结构理解的药物研发与分子设计场景

近年来,大语言模型推动了分子科学的发展。然而,仅依赖分子序列难以捕捉复杂结构。分子具有两种互补结构视角:一是原子间的拓扑关系(图视角),二是分子的空间构型(图像视角)。为协同利用这两类视角,我们提出跨视图前缀(CROP)框架,通过高效融合多视图信息增强分子LLM的理解能力。CROP具备两大优势:(i) 高效性——将多个结构视图联合重采样为固定长度前缀,避免占用过多上下文长度,且易于扩展至更多视图;(ii) 有效性——利用LLM自编码的分子序列指导重采样过程,提升生成前缀的质量。具体包括:基于SMILES的引导重采样器与结构嵌入门控模块,用于视图转换与前缀生成。大量实验表明,CROP在分子描述生成、IUPAC命名预测和分子性质预测等任务中均表现更优。

原文摘要 · Abstract (English)

Recent advances in molecular science have been propelled significantly by large language models (LLMs). However, their effectiveness is limited when relying solely on molecular sequences, which fail to capture the complex structures of molecules. Beyond sequence representation, molecules exhibit two complementary structural views: the first focuses on the topological relationships between atoms, as exemplified by the graph view; and the second emphasizes the spatial configuration of molecules, as represented by the image view. The two types of views provide unique insights into molecular structures. To leverage these views collaboratively, we propose the CROss-view Prefixes (CROP) to enhance LLMs' molecular understanding through efficient multi-view integration. CROP possesses two advantages: (i) efficiency: by jointly resampling multiple structural views into fixed-length prefixes, it avoids excessive consumption of the LLM's limited context length and allows easy expansion to more views; (ii) effectiveness: by utilizing the LLM's self-encoded molecular sequences to guide the resampling process, it boosts the quality of the generated prefixes. Specifically, our framework features a carefully designed SMILES Guided Resampler for view resampling, and a Structural Embedding Gate for converting the resulting embeddings into LLM's prefixes. Extensive experiments demonstrate the superiority of CROP in tasks including molecule captioning, IUPAC name prediction and molecule property prediction.

分子建模多视图学习语言模型结构理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。