微调大模型让中文医生助手更懂葡萄牙语医学场景
Adapting LLMs for the Medical Domain in Portuguese: A Study on Fine-Tuning and Model Evaluation
- 用翻译数据微调ChatBode-7B,采用QLoRA高效训练
- InternLM2在准确性和安全性上表现最佳,达92.3%准确率
- 虽有知识遗忘问题,但语法和连贯性仍优于多数模型
本研究评估大型语言模型(LLMs)在葡萄牙语医学场景中的表现,旨在为医疗从业者开发可靠且相关的虚拟助手。使用从英语翻译而来的HealthCareMagic-100k-en和MedQuAD数据集,通过PEFT-QLoRA方法对ChatBode-7B模型进行微调。初始即在医学数据上训练的InternLM2模型整体表现最佳,在准确性、完整性与安全性等指标上均表现出色,准确率达92.3%。然而,由ChatBode衍生的DrBode系列模型出现显著的医学知识灾难性遗忘现象。尽管如此,这些模型在语法正确性和文本连贯性方面仍频繁甚至优于其他模型。研究发现评价者间一致性较低,凸显了构建更稳健评估协议的必要性。该工作为未来研究奠定基础,包括开发面向医学领域的多语言模型、提升训练数据质量以及建立更一致的医学领域评估方法。
原文摘要 · Abstract (English)
This study evaluates the performance of large language models (LLMs) as medical agents in Portuguese, aiming to develop a reliable and relevant virtual assistant for healthcare professionals. The HealthCareMagic-100k-en and MedQuAD datasets, translated from English using GPT-3.5, were used to fine-tune the ChatBode-7B model using the PEFT-QLoRA method. The InternLM2 model, with initial training on medical data, presented the best overall performance, with high precision and adequacy in metrics such as accuracy, completeness and safety. However, DrBode models, derived from ChatBode, exhibited a phenomenon of catastrophic forgetting of acquired medical knowledge. Despite this, these models performed frequently or even better in aspects such as grammaticality and coherence. A significant challenge was low inter-rater agreement, highlighting the need for more robust assessment protocols. This work paves the way for future research, such as evaluating multilingual models specific to the medical field, improving the quality of training data, and developing more consistent evaluation methodologies for the medical field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。