让视觉语言模型更懂分子图像,提升药物设计理解能力
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding

- 引入分子拓扑模块,将化学键信息注入视觉特征
- 在多个化学图像任务中超越现有模型表现
- 适合药物发现与分子结构分析的研究者使用
利用分子大语言模型(LLMs)作为统一框架理解分子结构与功能,正成为分子设计和药物发现的新趋势。然而,这些模型难以充分捕捉分子结构的视觉表征,限制了其潜力。现有分子视觉语言模型(VLMs)虽有前景,但在结构对齐和拓扑建模方面仍存在不足。为此,我们提出 MolSight,一种图感知的视觉语言模型框架,通过分子拓扑模块将化学键邻接信息注入视觉标记,并通过分子定位模块对齐视觉特征与化学符号语义。实验表明,MolSight 在多个化学视觉理解任务中显著优于现有 VLMs、分子 LLMs 及专用模型,在复杂化学场景下实现了分子图像推理的新水平。
原文摘要 · Abstract (English)
Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a new trend in tasks such as molecular design and drug discovery. However, these models struggle to fully capture the visual representation of molecular structures, limiting their potential. While existing molecular vision-language models (VLMs) show promise, they still face challenges in structural alignment and lack the necessary topological modeling for accurate molecular understanding. To address this, we propose MolSight, a graph-aware vision-language model framework designed to enhance the understanding of molecular images by VLMs. MolSight integrates a Molecular Topology Module to inject chemical-bond adjacency information into vision tokens, and a Molecular Grounding Module to align visual features with chemical symbolic semantics. Our experiments demonstrate that MolSight significantly outperforms existing VLMs, molecular LLMs, and task-specific models across multiple chemical visual understanding tasks, achieving a new level of molecular image reasoning in complex chemical scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。