arXiv:2509.10095cs.CL2025-09被引 7

用阿拉伯语微调大模型,让医疗聊天更准更实用。

Arabic Large Language Models for Medical Text Generation

  • 用社交媒体真实医患对话微调LLM,适配多阿拉伯方言。
  • 微调后的Mistral-7B在关键指标上达68.5%~69.08%。
  • 适合资源有限地区,尤其需要多语言医疗AI的场景。

高效医院管理系统对缓解全球范围内的拥挤、资源短缺和紧急医疗可及性差等问题至关重要。现有方法常无法提供准确、实时的医疗建议,尤其针对非标准输入和低资源语言。为此,本研究提出一种面向阿拉伯语医疗文本生成的大语言模型微调方法,旨在通过用户输入为患者提供准确的诊断、药物推荐与治疗方案。研究收集了来自社交平台的真实医患对话数据集,涵盖患者主诉与医生建议,并完成清洗与预处理以应对多种阿拉伯方言。通过微调Mistral-7B-Instruct-v0.2、LLaMA-2-7B和GPT-2 Medium等先进生成模型,提升了系统生成可靠医疗文本的能力。评估结果显示,微调后的Mistral-7B在精确率、召回率和F1分数上的平均BERTScore分别为68.5%、69.08%和68.5%。对比基准测试与定性分析验证了系统对非正式输入生成连贯、相关医疗回复的能力。该研究展示了生成式AI在提升医院管理系统的潜力,为全球医疗挑战,特别是在语言文化多元环境中,提供了可扩展、可适应的解决方案。

原文摘要 · Abstract (English)

Efficient hospital management systems (HMS) are critical worldwide to address challenges such as overcrowding, limited resources, and poor availability of urgent health care. Existing methods often lack the ability to provide accurate, real-time medical advice, particularly for irregular inputs and underrepresented languages. To overcome these limitations, this study proposes an approach that fine-tunes large language models (LLMs) for Arabic medical text generation. The system is designed to assist patients by providing accurate medical advice, diagnoses, drug recommendations, and treatment plans based on user input. The research methodology required the collection of a unique dataset from social media platforms, capturing real-world medical conversations between patients and doctors. The dataset, which includes patient complaints together with medical advice, was properly cleaned and preprocessed to account for multiple Arabic dialects. Fine-tuning state-of-the-art generative models, such as Mistral-7B-Instruct-v0.2, LLaMA-2-7B, and GPT-2 Medium, optimized the system's ability to generate reliable medical text. Results from evaluations indicate that the fine-tuned Mistral-7B model outperformed the other models, achieving average BERT (Bidirectional Encoder Representations from Transformers) Score values in precision, recall, and F1-scores of 68.5\%, 69.08\%, and 68.5\%, respectively. Comparative benchmarking and qualitative assessments validate the system's ability to produce coherent and relevant medical replies to informal input. This study highlights the potential of generative artificial intelligence (AI) in advancing HMS, offering a scalable and adaptable solution for global healthcare challenges, especially in linguistically and culturally diverse environments.

医疗AI阿拉伯语大模型微调生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。