用规则引导自动化生成16万组分子结构描述,精度超98%。
A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method
- 基于规则解析IUPAC名生成结构元数据,驱动大模型生成精准描述
- 构建16.3万对分子-语言数据,人工+AI验证精度达98.6%
- 适合化学大模型训练、分子属性预测等需要结构语义对齐的任务
分子功能由其结构决定,准确对齐分子结构与自然语言对大语言模型开展下游化学任务推理至关重要。然而,人工标注成本高昂,难以构建大规模高质量结构描述数据集。本文提出一种全自动化标注框架,可规模化生成保留完整结构信息的分子描述。该方法在规则化化学命名解析基础上,将IUPAC名称转化为富含结构信息的XML元数据,并以此指导大语言模型生成自然语言描述。基于此框架,我们构建了约16.3万组分子-描述对。通过结合大模型与专家人工评估的严格验证协议,在2,000个分子子集上实现98.6%的描述精确率。所提框架可广泛服务于依赖结构描述的化学任务,生成的数据集为分子-语言对齐提供了可靠基础。源代码与数据集分别发布于https://github.com/TheLuoFengLab/MolLangData和https://huggingface.co/datasets/ChemFM/MolLangData。
原文摘要 · Abstract (English)
Molecular function is largely determined by structure. Accurately aligning molecular structure with natural language is therefore essential for enabling large language models (LLMs) to reason about downstream chemical tasks. However, the substantial cost of human annotation makes it infeasible to construct large-scale, high-quality datasets of structure-grounded descriptions. In this work, we propose a fully automated annotation framework for generating precise molecular descriptions that preserve complete structural details at scale. Our approach builds upon and extends a rule-based chemical nomenclature parser to interpret IUPAC names and construct enriched, structural XML metadata that explicitly encodes molecular structure. This metadata is then used to guide LLMs in producing accurate natural-language descriptions. Using this framework, we curate a large-scale dataset of approximately $163$k molecule--description pairs. A rigorous validation protocol combining LLM-based and expert human evaluation on a subset of $2,000$ molecules demonstrates a high description precision of $98.6$%. The proposed annotation framework is readily beneficial to broader chemical tasks that rely on structural descriptions, with the resulting dataset providing a reliable foundation for molecule--language alignment. The source code and dataset are hosted at https://github.com/TheLuoFengLab/MolLangData and https://huggingface.co/datasets/ChemFM/MolLangData, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。