对比多种大模型在生物医学任务中的表现,发现开源模型可媲美闭源模型。
Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
- 测试多款开源与闭源大模型在文本与图像任务上的表现
- 无单一模型在所有任务上最优,不同模型各有专长
- 开源模型在推理速度与隐私保护上更具优势
本文全面评估了多种成本效益高的大型语言模型(LLMs)在涵盖文本与图像模态的多样化生物医学任务中的表现。我们对一系列闭源与开源的LLMs进行了测试,任务包括生物医学文本分类与生成、问答以及多模态图像处理。实验结果表明,不存在一种能在所有任务中持续胜出的LLMs;相反,不同模型在不同任务中表现出色。尽管部分闭源模型在特定任务上表现优异,但其开源版本往往能达到相近甚至更优的结果,同时具备更快的推理速度和更强的隐私保护能力。研究为针对特定生物医学应用选择最优模型提供了重要参考。
原文摘要 · Abstract (English)
This paper presents a comprehensive evaluation of cost-efficient Large Language Models (LLMs) for diverse biomedical tasks spanning both text and image modalities. We evaluated a range of closed-source and open-source LLMs on tasks such as biomedical text classification and generation, question answering, and multimodal image processing. Our experimental findings indicate that there is no single LLM that can consistently outperform others across all tasks. Instead, different LLMs excel in different tasks. While some closed-source LLMs demonstrate strong performance on specific tasks, their open-source counterparts achieve comparable results (sometimes even better), with additional benefits like faster inference and enhanced privacy. Our experimental results offer valuable insights for selecting models that are optimally suited for specific biomedical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。