arXiv:2409.00084cs.CLcs.AI2024-09被引 17

评测多个大模型在胃肠病学领域的表现,发现闭源模型更优,视觉模型难用。

Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models

  • 对比12个主流大模型在胃肠病考题上的表现,涵盖闭源与开源版本
  • 闭源模型GPT-4o和Claude3.5准确率超73%,开源中Llama3.1-405b达64%
  • 量化模型性能接近全精度版,但图像理解仍依赖人工描述

本研究评估了大型语言模型(LLMs)和视觉语言模型(VLMs)在胃肠病学领域的医学推理能力。采用300道胃肠病学执业考试风格的多选题,其中138道含图像,系统评估模型配置、参数及提示工程策略对性能的影响。随后,在不同接口(网页端与API)、计算环境(云端与本地)和模型精度(量化与否)下,测试了包括GPT(3.5, 4, 4o, 4omini)、Claude(3, 3.5)、Gemini(1.0)、Mistral、Llama(2, 3, 3.1)、Mixtral和Phi(3)在内的专有与开源模型。最终通过半自动流程评估准确率。结果显示,闭源模型中GPT-4o(73.7%)和Claude3.5-Sonnet(74.0%)表现最佳,优于顶级开源模型:Llama3.1-405b(64%)、Llama3.1-70b(58.3%)和Mixtral-8x7b(54.3%)。在量化模型中,6比特量化的Phi3-14b(48.7%)表现最优。量化模型得分与全精度的Llama2-7b、Llama2-13b和Gemma2-9b相当。值得注意的是,当提供图像时,VLM在含图题目上的表现未提升,甚至在使用LLM生成的图像描述时下降;而人类编写的图像描述可使准确率提高10%。结论:尽管大模型在医学推理中展现出强大零样本能力,但视觉数据整合仍是视觉语言模型的挑战。有效部署需精心选择模型配置,用户应权衡闭源模型的高性能或开源模型的灵活性。

原文摘要 · Abstract (English)

Background and Aims: This study evaluates the medical reasoning performance of large language models (LLMs) and vision language models (VLMs) in gastroenterology. Methods: We used 300 gastroenterology board exam-style multiple-choice questions, 138 of which contain images to systematically assess the impact of model configurations and parameters and prompt engineering strategies utilizing GPT-3.5. Next, we assessed the performance of proprietary and open-source LLMs (versions), including GPT (3.5, 4, 4o, 4omini), Claude (3, 3.5), Gemini (1.0), Mistral, Llama (2, 3, 3.1), Mixtral, and Phi (3), across different interfaces (web and API), computing environments (cloud and local), and model precisions (with and without quantization). Finally, we assessed accuracy using a semiautomated pipeline. Results: Among the proprietary models, GPT-4o (73.7%) and Claude3.5-Sonnet (74.0%) achieved the highest accuracy, outperforming the top open-source models: Llama3.1-405b (64%), Llama3.1-70b (58.3%), and Mixtral-8x7b (54.3%). Among the quantized open-source models, the 6-bit quantized Phi3-14b (48.7%) performed best. The scores of the quantized models were comparable to those of the full-precision models Llama2-7b, Llama2--13b, and Gemma2-9b. Notably, VLM performance on image-containing questions did not improve when the images were provided and worsened when LLM-generated captions were provided. In contrast, a 10% increase in accuracy was observed when images were accompanied by human-crafted image descriptions. Conclusion: In conclusion, while LLMs exhibit robust zero-shot performance in medical reasoning, the integration of visual data remains a challenge for VLMs. Effective deployment involves carefully determining optimal model configurations, encouraging users to consider either the high performance of proprietary models or the flexible adaptability of open-source models.

大模型评测医学AI视觉语言模型量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。