首个分子属性预测语言增强基准,测试大模型如何提升毒性等预测性能。
MolCap-Arena: A Comprehensive Captioning Benchmark on Language-Enhanced Molecular Property Prediction
- 构建基于对抗评分的基准,评估20多个大模型在分子属性预测中的表现。
- 发现大模型提取的知识可显著提升现有分子表征效果,但效果因模型和数据而异。
- 适合关注分子生成、AI制药及大模型应用的研究者参考。
将生物分子建模与自然语言信息结合,特别是利用大语言模型(LLMs),已成为一个有前景的跨学科研究方向。经过大量科学文献训练的LLMs在理解与推理生物分子方面展现出巨大潜力,能提供丰富的上下文和领域知识。然而,这些由LLM驱动的洞察对复杂预测任务(如毒性预测)的提升程度尚不明确,且从中提取相关知识的能力也未被充分了解。本研究提出分子描述竞技场(MolCap-Arena):首个面向语言增强分子属性预测的综合性基准。我们评估了超过二十个大模型,包括通用和领域专用的分子描述生成器,在多种预测任务上的表现。为此,引入了一种新颖的基于对抗评分的评级系统。研究结果证实,从LLM中提取的知识能够提升最先进的分子表征性能,但存在显著的模型、提示和数据集依赖性差异。代码、资源和数据已公开于github.com/Genentech/molcap-arena。
原文摘要 · Abstract (English)
Bridging biomolecular modeling with natural language information, particularly through large language models (LLMs), has recently emerged as a promising interdisciplinary research area. LLMs, having been trained on large corpora of scientific documents, demonstrate significant potential in understanding and reasoning about biomolecules by providing enriched contextual and domain knowledge. However, the extent to which LLM-driven insights can improve performance on complex predictive tasks (e.g., toxicity) remains unclear. Further, the extent to which relevant knowledge can be extracted from LLMs also remains unknown. In this study, we present Molecule Caption Arena: the first comprehensive benchmark of LLM-augmented molecular property prediction. We evaluate over twenty LLMs, including both general-purpose and domain-specific molecule captioners, across diverse prediction tasks. To this goal, we introduce a novel, battle-based rating system. Our findings confirm the ability of LLM-extracted knowledge to enhance state-of-the-art molecular representations, with notable model-, prompt-, and dataset-specific variations. Code, resources, and data are available at github.com/Genentech/molcap-arena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。