构建分子语言交互评测基准,揭示当前AI在分子识别与生成上的严重不足。
MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation
- 基于专家标注和自动化工具构建三类分子任务,确保输出清晰可靠。
- 最强模型在识别/编辑任务上准确率仅86.2%和85.5%,生成任务更低至43.0%。
- 适合关注化学AI、分子生成与语言接口研究的研究者使用。
精确的分子识别、编辑与生成是化学家与AI系统完成各类化学任务的基础前提。我们提出MolLangBench,一个全面的基准,用于评估分子-语言接口的核心任务:语言提示下的分子结构识别、编辑与生成。为确保高质量、无歧义且确定性的输出,我们利用自动化化学信息学工具构建识别任务,并通过严格专家标注与验证来整理编辑与生成任务。MolLangBench支持评估对接不同分子表示(如线性字符串、分子图像、分子图)的语言模型。对最先进模型的评估显示显著局限:最强模型(GPT-5)在识别和编辑任务上分别达到86.2%和85.5%的准确率,而人类看来直观简单的任务,其生成任务准确率仅43.0%。这些结果凸显当前AI系统在处理基础分子识别与操作任务时的不足。我们希望MolLangBench能推动更高效可靠的化学应用AI研究。数据集与代码可访问于https://huggingface.co/datasets/ChemFM/MolLangBench 和 https://github.com/TheLuoFengLab/MolLangBench。
原文摘要 · Abstract (English)
Precise recognition, editing, and generation of molecules are essential prerequisites for both chemists and AI systems tackling various chemical tasks. We present MolLangBench, a comprehensive benchmark designed to evaluate fundamental molecule-language interface tasks: language-prompted molecular structure recognition, editing, and generation. To ensure high-quality, unambiguous, and deterministic outputs, we construct the recognition tasks using automated cheminformatics tools, and curate editing and generation tasks through rigorous expert annotation and validation. MolLangBench supports the evaluation of models that interface language with different molecular representations, including linear strings, molecular images, and molecular graphs. Evaluations of state-of-the-art models reveal significant limitations: the strongest model (GPT-5) achieves $86.2\%$ and $85.5\%$ accuracy on recognition and editing tasks, which are intuitively simple for humans, and performs even worse on the generation task, reaching only $43.0\%$ accuracy. These results highlight the shortcomings of current AI systems in handling even preliminary molecular recognition and manipulation tasks. We hope MolLangBench will catalyze further research toward more effective and reliable AI systems for chemical applications.The dataset and code can be accessed at https://huggingface.co/datasets/ChemFM/MolLangBench and https://github.com/TheLuoFengLab/MolLangBench, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。