首个评估大模型开放域分子生成能力的基准,突破单一答案限制。
Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation
- 构建一语多解的分子生成评测框架,模拟真实设计场景。
- 31个大模型测试显示,Llama3.1-8B在新数据集上超越GPT-4o等模型。
- 适合药物发现、化学AI研究者,推动语言驱动分子设计发展。
大语言模型在自然语言驱动的分子发现中展现出巨大潜力,但现有数据集和评测主要基于一对一映射,仅衡量模型对预设答案的检索能力,而非创造性生成多样且有效的分子候选物的能力。为填补这一关键空白,我们提出首个开放域自然语言驱动分子生成评测基准Speak-to-Structure (S^2-Bench),专门针对一语多解关系,挑战模型的真实分子理解与开放式生成能力。该基准包含三个核心任务:分子编辑(MolEdit)、分子优化(MolOpt)和定制化分子生成(MolCustom),分别探测分子发现的不同方面。我们还引入OpenMolIns大规模指令微调数据集,使Llama3.1-8B在S^2-Bench上超越GPT-4o和Claude-3.5等最强模型。对31个大模型的全面评估将关注点从简单模式回忆转向真实的分子设计,为自然语言驱动的分子发现铺平道路。代码与数据集已通过GitHub(https://github.com/phenixace/S2-TOMG-Bench)和HuggingFace数据集平台公开。
原文摘要 · Abstract (English)
Recently, Large Language Models (LLMs) have demonstrated great potential in natural language-driven molecule discovery. However, existing datasets and benchmarks for molecule-text alignment are predominantly built on one-to-one mappings, measuring LLMs' ability to retrieve a single, pre-defined answer, rather than their creative potential to generate diverse, yet equally valid, molecular candidates. To address this critical gap, we propose Speak-to-Structure (S^2-Bench), the first benchmark to evaluate LLMs in open-domain natural language-driven molecule generation. S^2-Bench is specifically designed for one-to-many relationships, challenging LLMs to exhibit genuine molecular understanding and open-ended generation capabilities. Our benchmark includes three key tasks: molecule editing (MolEdit), molecule optimization (MolOpt), and customized molecule generation (MolCustom), each probing a different aspect of molecule discovery. We also introduce OpenMolIns, a large-scale instruction tuning dataset that enables Llama3.1-8B to surpass the most powerful LLMs like GPT-4o and Claude-3.5 on S^2-Bench. Our comprehensive evaluation of 31 LLMs shifts the focus from simple pattern recall to realistic molecular design, paving the way for more capable LLMs in natural language-driven molecule discovery. Our codes and datasets are fully accessible through the Github Repository: https://github.com/phenixace/S2-TOMG-Bench and Huggingface Datasets: https://huggingface.co/datasets/phenixace/S2-TOMG-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。