测试大模型在巴西葡语医学考试中的零样本表现,发现部分模型接近人类水平。
Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam
- 用巴西葡语医学考题测试6个LLM和4个MLLM的零样本能力。
- Claude-3.5-Sonnet等模型准确率接近人类,但图像理解仍有差距。
- 揭示非英语医疗AI的性能鸿沟,呼吁加强多语言训练数据。
人工智能在医疗领域展现出提升诊断精度、优化流程与个性化治疗的潜力。大型语言模型(LLMs)和多模态大语言模型(MLLMs)在自然语言处理与医疗应用中取得显著进展,但现有评估主要集中在英语,可能导致跨语言性能偏差。本研究考察了六种LLM(GPT-4.0 Turbo、LLaMA-3-8B、LLaMA-3-70B、Mixtral 8x7B Instruct、Titan Text G1-Express、Command R+)和四种MLLM(Claude-3.5-Sonnet、Claude-3-Opus、Claude-3-Sonnet、Claude-3-Haiku)在巴西葡语医学考试(来自圣保罗大学医学院临床医院,HCFMUSP,南美最大医疗中心)中的表现。模型性能与人类考生对比,分析准确率、处理时间与解释连贯性。结果显示,部分模型如Claude-3.5-Sonnet和Claude-3-Opus准确率接近人类,但在需图像理解的多模态问题上仍存在明显差距。研究强调语言差异,呼吁为非英语医疗AI进一步微调与数据增强。结论强调应在多种语言与临床场景下评估生成式AI,以确保医疗部署的公平性与可靠性。未来研究应探索更优训练方法、多模态推理能力及真实临床整合。
原文摘要 · Abstract (English)
Artificial intelligence (AI) has shown the potential to revolutionize healthcare by improving diagnostic accuracy, optimizing workflows, and personalizing treatment plans. Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have achieved notable advancements in natural language processing and medical applications. However, the evaluation of these models has focused predominantly on the English language, leading to potential biases in their performance across different languages. This study investigates the capability of six LLMs (GPT-4.0 Turbo, LLaMA-3-8B, LLaMA-3-70B, Mixtral 8x7B Instruct, Titan Text G1-Express, and Command R+) and four MLLMs (Claude-3.5-Sonnet, Claude-3-Opus, Claude-3-Sonnet, and Claude-3-Haiku) to answer questions written in Brazilian spoken portuguese from the medical residency entrance exam of the Hospital das Clínicas da Faculdade de Medicina da Universidade de São Paulo (HCFMUSP) - the largest health complex in South America. The performance of the models was benchmarked against human candidates, analyzing accuracy, processing time, and coherence of the generated explanations. The results show that while some models, particularly Claude-3.5-Sonnet and Claude-3-Opus, achieved accuracy levels comparable to human candidates, performance gaps persist, particularly in multimodal questions requiring image interpretation. Furthermore, the study highlights language disparities, emphasizing the need for further fine-tuning and data set augmentation for non-English medical AI applications. Our findings reinforce the importance of evaluating generative AI in various linguistic and clinical settings to ensure a fair and reliable deployment in healthcare. Future research should explore improved training methodologies, improved multimodal reasoning, and real-world clinical integration of AI-driven medical assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。