构建30个生物信息任务的提示评估框架,测试大模型在无微调下的真实能力。
Benchmarking Large Language Models on Multiple Tasks in Bioinformatics NLP with Prompting
- 设计基于提示的评测框架,覆盖蛋白质、药物等30类生物任务。
- GPT-4o和Llama-3.1-70b在零样本下表现最佳,准确率超60%。
- 适合生物信息学研究者与AI开发者参考,推动模型优化方向。
大语言模型(LLMs)已成为解决生物学问题的重要工具,相较于传统方法,在准确性和适应性方面均有提升。然而,现有基准难以有效评估模型在多样化任务中的表现。本文提出一个全面的提示驱动评测框架Bio-benchmark,涵盖蛋白质、RNA、药物、电子健康记录及中医等领域的30项关键生物信息任务。我们利用该框架,在不进行微调的情况下,以零样本和少样本思维链(CoT)设置评估了六款主流大模型,包括GPT-4o和Llama-3.1-70b,揭示其内在能力。为提升评估效率,我们提出了BioFinder工具,可从模型输出中提取答案,相比现有方法准确率提升约30%。评测结果明确了当前大模型适用的生物任务,并指出了需改进的具体领域。此外,我们提出了针对性的提示工程策略以优化性能。基于这些发现,我们为开发更适用于生物应用的大模型提供了建议。本工作提供了一个全面的评估框架与可靠工具,支持大模型在生物信息学中的落地应用。
原文摘要 · Abstract (English)
Large language models (LLMs) have become important tools in solving biological problems, offering improvements in accuracy and adaptability over conventional methods. Several benchmarks have been proposed to evaluate the performance of these LLMs. However, current benchmarks can hardly evaluate the performance of these models across diverse tasks effectively. In this paper, we introduce a comprehensive prompting-based benchmarking framework, termed Bio-benchmark, which includes 30 key bioinformatics tasks covering areas such as proteins, RNA, drugs, electronic health records, and traditional Chinese medicine. Using this benchmark, we evaluate six mainstream LLMs, including GPT-4o and Llama-3.1-70b, etc., using 0-shot and few-shot Chain-of-Thought (CoT) settings without fine-tuning to reveal their intrinsic capabilities. To improve the efficiency of our evaluations, we demonstrate BioFinder, a new tool for extracting answers from LLM responses, which increases extraction accuracy by round 30% compared to existing methods. Our benchmark results show the biological tasks suitable for current LLMs and identify specific areas requiring enhancement. Furthermore, we propose targeted prompt engineering strategies for optimizing LLM performance in these contexts. Based on these findings, we provide recommendations for the development of more robust LLMs tailored for various biological applications. This work offers a comprehensive evaluation framework and robust tools to support the application of LLMs in bioinformatics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。