首个面向波斯语医学问答的预训练大模型,支持长答案生成
BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining
- 基于1万+文献构建波斯语医学数据集,设计专用评估任务
- 在5231个医学问答上达ROUGE-L 29.99,优于GPT-4 1.0
- 适合波斯语生物医学信息提取与智能问答研究者使用
大型语言模型(LLMs)在生命科学领域日益受到关注,因其能够建模、提取和应用复杂的生物信息。本文首先构建了包含超10,000篇科学论文、教科书及医疗网站内容的BIOPARS-BENCH数据集,并引入BioParsQA,包含5,231个波斯语医学问答对。提出BioPars模型,用于评估模型在获取专业知识、解释合成知识及提供证据三方面的能力。对比ChatGPT、Llama与Galactica发现,其虽具备知识记忆与检索能力,但在解决复杂真实问题和细粒度推理上仍存不足。本研究首次将LLM应用于波斯语医学问答,尤其长答案生成。在四个医学QA数据集上的评估显示,该模型在多项指标上表现优异:在BioParsQA上取得ROUGE-L 29.99,高于GPT-4 1.0;BERTScore达90.87(使用MMR方法);MoverScore=60.43,BLEURT=50.78,均优于其他三模型。相关资源已开源于GitHub。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently gained attention in the life sciences due to their capacity to model, extract, and apply complex biological information. Beyond their classical use as chatbots, these systems are increasingly used for complex analysis and problem-solving in specialized fields, including bioinformatics. First, we introduce BIOPARS-BENCH, a dataset from over 10,000 scientific articles, textbooks, and medical websites. BioParsQA was also introduced to evaluate the proposed model, which consists of 5,231 Persian medical questions and answers. This study then introduces BioPars, a simple but accurate measure designed to assess LLMs for three main abilities: acquiring subject-specific knowledge, interpreting and synthesizing such knowledge, and demonstrating proper evidence. Comparing ChatGPT, Llama, and Galactica, our study highlights their ability to remember and retrieve learned knowledge but also reveals shortcomings in addressing higher-level, real-world questions and fine-grained inferences. These findings indicate the need for further fine-tuning to address the capabilities of LLM in bioinformatics tasks. To our knowledge, BioPars is the first application of LLM in Persian medical QA, especially for generating long answers. Evaluation of four selected medical QA datasets shows that BioPars has achieved remarkable results compared to comparative approaches. The model on BioParsQA achieved a ROUGE-L score of 29.99, which is an improvement over GPT-4 1.0. The model achieved a BERTScore of 90.87 with the MMR method. The MoverScore and BLEURT values were also higher in this model than the other three models. In addition, the reported scores for the model are MoverScore=60.43 and BLEURT=50.78. BioPars is an ongoing project and all resources related to its development will be made available via the following GitHub repository: https://github.com/amirap80/BioPars.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。