首个专为分子编程任务设计的评测基准,检验大模型生成可执行化学代码能力。
MolViBench: Evaluating LLMs on Molecular Vibe Coding

- 构建跨五认知层级的分子编程任务集,涵盖从单接口调用到全流程药物筛选
- 提出多层评估框架,结合类型比对与语法树分析,精准判断代码可执行性与化学正确性
- 适用于药物研发人员、AI+化学交叉研究者,助力大模型在分子工作流中的落地
分子共鸣编程(Molecular Vibe Coding)是一种新范式,让化学家通过大语言模型(LLM)生成可执行程序完成分子任务,无需预设工具链,可灵活定义复杂定制化流程。该任务要求模型兼具编程能力、分子理解与领域推理能力,但现有基准存在断层:通用编码基准如HumanEval无需化学知识,而化学专用基准如S^2-Bench和ChemCoTBench仅评估知识记忆或性质预测,不涉及可执行代码生成。为此,我们提出MolViBench,首个专为分子共鸣编程设计的基准。其包含358个精选任务,覆盖五个认知层级,从单接口调用到端到端虚拟筛选流程设计,涵盖12个真实世界药物发现流程。为严谨评估生成代码,我们设计多层次评估框架,融合类型感知输出对比与基于抽象语法树(AST)的API语义回退分析,联合衡量代码可执行性与化学正确性。我们系统评估了9个前沿编码大模型,并对比三种真实世界的分子共鸣编程范式,提供了一个实用且细粒度的测试平台,用于诊断大模型在人工智能加速分子发现中的编程能力。
原文摘要 · Abstract (English)
Molecular Vibe Coding, a paradigm where chemists interact with LLMs to generate executable programs for molecular tasks, has emerged as a flexible alternative to chemical agents with predefined tools, enabling chemists to express arbitrarily complex, customized workflows. Unlike general coding tasks, molecular coding imposes a distinctive challenge that LLMs should jointly equip programming, molecular understanding, and domain-specific reasoning capabilities. However, existing benchmarks remain disconnected. General code generation benchmarks such as HumanEval and SWE-bench require no chemistry knowledge, while chemistry-focused benchmarks such as S^2-Bench and ChemCoTBench evaluate knowledge recall or property prediction rather than executable code generation. To bridge this gap, we introduce MolViBench, the first benchmark tailored for Molecular Vibe Coding. MolViBench comprises 358 curated tasks across five cognitive levels, ranging from single-API recall to end-to-end virtual screening pipeline design, spanning 12 real-world drug discovery workflows. To rigorously assess generated code, we also propose a multi-layered evaluation framework that combines type-aware output comparison and AST-based API-semantic fallback analysis, which jointly measures executability and chemical correctness. We systematically evaluate 9 frontier coding LLMs and compare three real-world Molecular Vibe Coding paradigms, providing a practical and fine-grained testbed for diagnosing LLMs' coding capabilities in AI-accelerated molecular discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。