用在线医学数据提升小模型的波斯语医疗能力
Leveraging Online Data to Enhance Medical Knowledge in a Small Persian Language Model
- 从网页爬取并构建首个波斯语医学问答数据集
- 模型通过微调在伊朗医考中达标,翻译MMLU得分提升2.67%
- 适合资源有限环境下开发波斯语医疗AI
语言模型的快速发展展示了人工智能在医疗领域的潜力。然而,小型语言模型在低资源语言如波斯语的专科领域表现不佳。尽管存在大量波斯语医学网站,但此前缺乏经过整理的数据集。本研究首次构建了一个包含20,000对医生-患者问答及9000万词元中60%爬取的医学杂志语料库的新数据集。采用参数高效微调方法,提升了基准模型aya-expanse-8b的医学知识。基准测试显示,微调后模型在医学问答任务中准确率提升,并成功通过2023年9月的伊朗基础医学科学入学考试(IBSEE),而原模型未通过。此外,该模型使波斯语翻译版MMLU平均得分提升2.67%。本工作证明了利用公开在线数据增强小型语言模型在医学领域的能力,为资源受限环境下的波斯语医疗AI提供了新方案。未来可探索多模态输入以进一步提升性能。
原文摘要 · Abstract (English)
The rapid advancement of language models has demonstrated the potential of artificial intelligence in the healthcare industry. However, small language models struggle with specialized domains in low-resource languages like Persian. While numerous medical-domain websites exist in Persian, no curated dataset or corpus has been available making ours the first of its kind. This study introduces a newly curated dataset comprising 20k doctor-patient Q\&A pairs and 60\% of a 90-million-token crawled corpus from medical magazines. Using a parameter-efficient fine-tuning approach, we enhanced the medical knowledge of the baseline model, aya-expanse-8b. Benchmark evaluations demonstrate that the fine-tuned model achieves improved accuracy in medical question answering and successfully passed the Iranian Basic Medical Science Entrance Exam (IBSEE) in September 2023, which the baseline model did not. Additionally, the fine-tuned model improved Persian-translated MMLU accuracy by an average of 2.67\%. This work highlights the potential of leveraging open-access online data to enrich small language models in medical fields, providing a novel solution for Persian medical AI applications suitable for resource-constrained environments. Future research could explore multimodal input to further enhance performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。