构建首个分子功能基团级属性推理数据集,助力大模型理解化学结构与性质关系。
FGBench: A Dataset and Benchmark for Molecular Property Reasoning at Functional Group-Level in Large Language Models
- 基于62.5万条数据,标注分子中功能基团位置与属性关系
- 涵盖245种功能基团的单个影响、交互作用及分子比较任务
- 揭示当前大模型在基团级推理能力不足,推动可解释性研究
大型语言模型(LLMs)在化学领域受到广泛关注。然而,现有数据集多聚焦于分子级属性预测,忽视了细粒度功能基团(FG)信息的作用。引入功能基团级数据可提供关键先验知识,建立分子结构与文本描述间的联系,有助于构建更具可解释性和结构感知能力的模型,用于分子相关任务推理。此外,模型可从中学习特定功能基团与分子性质间的隐含关系,推动分子设计与药物发现。本文提出FGBench,一个包含62.5万条功能基团级分子属性推理问题的数据集,其中功能基团被精确标注并定位在分子内,确保数据可互操作性,支持多模态应用。该数据集涵盖三类任务:(1)单个功能基团的影响,(2)多个功能基团的相互作用,(3)直接分子比较,覆盖245种不同功能基团。在7,000条精选数据上的基准测试显示,当前先进大模型在功能基团级推理上表现不佳,凸显提升其化学推理能力的必要性。我们预期,FGBench构建方法可为生成新问答对提供基础框架,帮助大模型更好理解分子结构-性质的精细关系。数据集与评估代码已开源:https://github.com/xuanliugit/FGBench。
原文摘要 · Abstract (English)
Large language models (LLMs) have gained significant attention in chemistry. However, most existing datasets center on molecular-level property prediction and overlook the role of fine-grained functional group (FG) information. Incorporating FG-level data can provide valuable prior knowledge that links molecular structures with textual descriptions, which can be used to build more interpretable, structure-aware LLMs for reasoning on molecule-related tasks. Moreover, LLMs can learn from such fine-grained information to uncover hidden relationships between specific functional groups and molecular properties, thereby advancing molecular design and drug discovery. Here, we introduce FGBench, a dataset comprising 625K molecular property reasoning problems with functional group information. Functional groups are precisely annotated and localized within the molecule, which ensures the dataset's interoperability thereby facilitating further multimodal applications. FGBench includes both regression and classification tasks on 245 different functional groups across three categories for molecular property reasoning: (1) single functional group impacts, (2) multiple functional group interactions, and (3) direct molecular comparisons. In the benchmark of state-of-the-art LLMs on 7K curated data, the results indicate that current LLMs struggle with FG-level property reasoning, highlighting the need to enhance reasoning capabilities in LLMs for chemistry tasks. We anticipate that the methodology employed in FGBench to construct datasets with functional group-level information will serve as a foundational framework for generating new question-answer pairs, enabling LLMs to better understand fine-grained molecular structure-property relationships. The dataset and evaluation code are available at https://github.com/xuanliugit/FGBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。