arXiv:2602.02742cs.LGcs.AI2026-02中稿 · ICLR被引 3

动态生成分子结构关键片段的注意力令牌,提升大模型对分子图的理解能力。

Entropy-Guided Dynamic Tokens for Graph-LLM Alignment in Molecular Understanding

  • 基于熵值动态生成与分子关键区域对齐的注意力令牌
  • 无需微调大模型主干,在多个基准上达到最优性能
  • 适合需要高效、通用分子理解的药物研发与化学智能场景

分子理解是推动科学发现的核心,但大语言模型(LLM)在理解分子图方面表现不佳。现有图-语言模型桥接方法多采用固定长度的静态令牌(源自视觉任务的Q-Former风格),忽略立体化学和子结构上下文,且通常需代价高昂的LLM主干微调,限制了效率与泛化性。本文提出EDT-Former:一种熵引导的动态令牌变换器,可生成与分子信息片段对齐的令牌,从而保留局部与全局结构特征。相比以往方法,EDT-Former可在不微调LLM主干(仅嵌入层除外)的前提下实现图编码器与语言模型的对齐,实现计算高效的微调,并在MoleculeQA、Molecule-oriented Mol-Instructions、TDC与MoleculeNet等属性预测基准上取得领先结果,验证了其在可扩展、通用的多模态分子理解中的有效性。

原文摘要 · Abstract (English)

Molecular understanding is central to advancing areas such as scientific discovery, yet Large Language Models (LLMs) struggle to understand molecular graphs effectively. Existing graph-LLM bridges often adapt the Q-Former-style connector with fixed-length static tokens, which is originally designed for vision tasks. These designs overlook stereochemistry and substructural context and typically require costly LLM-backbone fine-tuning, limiting efficiency and generalization. We introduce EDT-Former, an Entropy-guided Dynamic Token Transformer that generates tokens aligned with informative molecular patches, thereby preserving both local and global structural features for molecular graph understanding. Beyond prior approaches, EDT-Former enables alignment between frozen graph encoders and LLMs without tuning the LLM backbone (excluding the embedding layer), resulting in computationally efficient finetuning, and achieves stateof-the-art results on MoleculeQA, Molecule-oriented Mol-Instructions, and property prediction benchmarks (TDC, MoleculeNet), underscoring its effectiveness for scalable and generalizable multimodal molecular understanding

分子理解图神经网络动态令牌大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。