首个超分子化学大模型评测基准,助力智能设计分子结合系统。
SupraBench: A Benchmark for Supramolecular Chemistry

- 构建四类基础任务+视觉识别,系统评估大模型在超分子化学中的推理能力。
- 现有大模型在所有任务上仍有显著提升空间,尤其在格式严格输出上表现不佳。
- 适配超分子领域语料后,回归任务性能提升,但对输出格式要求敏感。
超分子化学研究非共价主客体组装,在多个领域具有重要应用。然而,设计主客体系统仍耗时费力,每个候选组合需数天的纯理论验证。尽管大语言模型(LLMs)在分子结合任务中表现出快速且高效的优势,目前尚无系统性基准用于评估其在超分子化学核心任务中的表现,如结合亲和力预测。为此,我们与领域专家合作,发布首个超分子化学基准——SupraBench,用于评估大模型在化学推理中的能力。具体设计了四项基础任务:结合亲和力预测、最优结合剂选择、溶剂识别及主客体描述,并额外增加一个基于视觉的分子识别辅助任务。同时发布了一个从欧洲文献库(Europe PMC)中提炼的1600万词的超分子化学语料库SupraPMC,以支持模型在该领域的适应。我们对多种开源与专有大模型进行了评测,发现所有任务均存在较大改进空间。在SupraPMC上进行领域预训练可有效提升分布内回归性能,但会牺牲对严格格式输出的准确性。此外,不同任务族的难度差异显著,暴露出当前模型在超分子化学推理中的特定缺陷。代码与数据集已公开于https://github.com/Tianyi-Billy-Ma/SupraBench。
原文摘要 · Abstract (English)
Supramolecular chemistry, which includes the study of non-covalent host-guest assemblies, has advanced various applications. However, designing host-guest systems remains time-consuming, requiring days of dry-lab verification per candidate pair. Although LLMs have emerged as a fast alternative with strong performance on molecular binding tasks, no benchmark currently systematically evaluates LLMs for host-guest reasoning across fundamental supramolecular chemistry tasks, e.g., binding affinity prediction. To this end, we collaborate with domain experts to release the first Supramolecular Benchmark, called SupraBench, to evaluate LLMs in chemistry reasoning. Specifically, we design four fundamental tasks, i.e., binding affinity prediction, top-binder selection, solvent identification, and host-guest description, plus an auxiliary vision-based task for molecular identification. We also release SupraPMC, a curated 16M-token corpus of Supramolecular chemistry articles distilled from Europe PMC, to support the adaptation to the supramolecular domain. We benchmark a broad range of open and proprietary LLMs and find that LLMs leave substantial headroom across all tasks. Domain adaptation pretraining over SupraPMC transfers cleanly to in-distribution regression but trades off against strict letter-format output. Moreover, the difficulty profile differs sharply across task families, revealing distinct failure modes that indicate specific gaps in current supramolecular chemistry reasoning. Our source codes and benchmark datasets are available at https://github.com/Tianyi-Billy-Ma/SupraBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。