Eir-8B是为泰语医疗场景打造的80亿参数大模型,显著提升诊疗效率与准确性。
Eir: Thai Medical Large Language Models
- 基于4个医疗基准数据集微调,融合多策略推理提升问答能力
- 在18项临床任务中超越GPT-4o超11%,泰语医疗表现领先商用模型10%以上
- 部署于医院内网,加密接口保障数据安全,适配医护人员与患者双端使用
我们提出Eir-8B,一个拥有80亿参数的大型语言模型,专为提升泰语医疗任务的准确性而设计。该模型致力于为医疗专业人员和患者提供清晰易懂的回答,从而提高诊断与治疗效率。通过人工评估确保模型符合医疗照护标准并提供无偏见答案。为保障数据安全,模型部署于医院内部网络,实现高安全性与快速处理。内部API连接采用加密与严格认证机制,防止数据泄露与未授权访问。我们在四个医疗基准(MedQA、MedMCQA、PubMedQA及MMLU医学子集)上评估多个80亿参数开源模型,并以此为基础构建Eir-8B。评估采用零样本、少样本、思维链推理及集成/自一致性投票等多重提问策略。结果显示,该模型在泰语环境下性能优于商用模型超过10%。此外,我们针对泰语临床场景开发了18项增强型测试,其表现超出GPT-4o超过11%。
原文摘要 · Abstract (English)
We present Eir-8B, a large language model with 8 billion parameters, specifically designed to enhance the accuracy of handling medical tasks in the Thai language. This model focuses on providing clear and easy-to-understand answers for both healthcare professionals and patients, thereby improving the efficiency of diagnosis and treatment processes. Human evaluation was conducted to ensure that the model adheres to care standards and provides unbiased answers. To prioritize data security, the model is deployed within the hospital's internal network, ensuring both high security and faster processing speeds. The internal API connection is secured with encryption and strict authentication measures to prevent data leaks and unauthorized access. We evaluated several open-source large language models with 8 billion parameters on four medical benchmarks: MedQA, MedMCQA, PubMedQA, and the medical subset of MMLU. The best-performing baselines were used to develop Eir-8B. Our evaluation employed multiple questioning strategies, including zero-shot, few-shot, chain-of-thought reasoning, and ensemble/self-consistency voting methods. Our model outperformed commercially available Thai-language large language models by more than 10%. In addition, we developed enhanced model testing tailored for clinical use in Thai across 18 clinical tasks, where our model exceeded GPT-4o performance by more than 11%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。