arXiv:2503.23668cs.AI2025-03被引 2

构建首个分子结构定位基准,让AI能精准定位分子中的具体化学部分。

MolGround: A Benchmark for Molecular Grounding

  • 基于NLP范式设计分子定位任务,建立可量化评估的参照体系。
  • 创建包含11.7万组问答对的最大分子理解基准数据集。
  • 多智能体系统超越GPT-4o,提升分子描述与分类任务性能。

当前分子理解方法主要关注人类感知的描述性层面,提供宽泛的主题级洞察,但对指称性层面——即分子概念与具体结构成分的关联——仍基本未被探索。为填补这一空白,我们提出一个分子定位基准,用于评估模型的指称能力。该基准遵循自然语言处理、化学信息学及分子科学的既有规范,展示了NLP技术在推动人工智能科学领域分子理解方面的潜力。此外,我们构建了迄今为止最大的分子理解基准,包含11.7万个问答对,并开发了一个多智能体定位原型系统作为概念验证。该系统性能优于现有模型(包括GPT-4o),其定位输出已集成至传统任务中,如分子描述生成和ATC(解剖学、治疗学、化学)分类任务,显著提升表现。

原文摘要 · Abstract (English)

Current molecular understanding approaches predominantly focus on the descriptive aspect of human perception, providing broad, topic-level insights. However, the referential aspect -- linking molecular concepts to specific structural components -- remains largely unexplored. To address this gap, we propose a molecular grounding benchmark designed to evaluate a model's referential abilities. We align molecular grounding with established conventions in NLP, cheminformatics, and molecular science, showcasing the potential of NLP techniques to advance molecular understanding within the AI for Science movement. Furthermore, we constructed the largest molecular understanding benchmark to date, comprising 117k QA pairs, and developed a multi-agent grounding prototype as proof of concept. This system outperforms existing models, including GPT-4o, and its grounding outputs have been integrated to enhance traditional tasks such as molecular captioning and ATC (Anatomical, Therapeutic, Chemical) classification.

分子理解指称定位AI for Science

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。