arXiv:2509.10108cs.CL2025-09中稿 · AICCSA 2025被引 1

用合成数据扩增10万条阿拉伯语医患对话,提升医疗聊天机器人性能。

Scaling Arabic Medical Chatbots Using Synthetic Data: Enhancing Generative AI with Synthetic Patient Records

  • 用ChatGPT-4o和Gemini生成8万条结构一致的合成问诊对。
  • 合成数据使模型在BERTScore上提升,幻觉减少37%以上。
  • 适合低资源医疗NLP研究者与阿拉伯语AI应用开发者。

阿拉伯语医疗聊天机器人的发展受限于高质量标注数据集的匮乏。此前研究基于社交媒体构建了2万条阿拉伯语医患对话数据集以微调大语言模型(LLMs),但模型的可扩展性和泛化能力仍不足。本研究提出一种可扩展的合成数据增强策略,将训练语料扩充至10万条。利用ChatGPT-4o和Gemini 2.5 Pro生成8万条语境相关且医学一致的合成问答对,其结构基于原始数据集。合成样本经语义过滤与人工验证后融入训练流程。我们微调了包括Mistral-7B和AraGPT2在内的五种LLM,使用BERTScore及专家评估进行性能测试。通过消融实验对比ChatGPT-4o与Gemini生成数据的效果,结果表明ChatGPT-4o生成数据在所有模型中均带来更高F1分数并显著减少幻觉。研究证明,合成数据增强是提升低资源医疗NLP领域专用语言模型的有效途径,为更包容、可扩展、精准的阿拉伯语医疗聊天机器人系统奠定基础。

原文摘要 · Abstract (English)

The development of medical chatbots in Arabic is significantly constrained by the scarcity of large-scale, high-quality annotated datasets. While prior efforts compiled a dataset of 20,000 Arabic patient-doctor interactions from social media to fine-tune large language models (LLMs), model scalability and generalization remained limited. In this study, we propose a scalable synthetic data augmentation strategy to expand the training corpus to 100,000 records. Using advanced generative AI systems ChatGPT-4o and Gemini 2.5 Pro we generated 80,000 contextually relevant and medically coherent synthetic question-answer pairs grounded in the structure of the original dataset. These synthetic samples were semantically filtered, manually validated, and integrated into the training pipeline. We fine-tuned five LLMs, including Mistral-7B and AraGPT2, and evaluated their performance using BERTScore metrics and expert-driven qualitative assessments. To further analyze the effectiveness of synthetic sources, we conducted an ablation study comparing ChatGPT-4o and Gemini-generated data independently. The results showed that ChatGPT-4o data consistently led to higher F1-scores and fewer hallucinations across all models. Overall, our findings demonstrate the viability of synthetic augmentation as a practical solution for enhancing domain-specific language models in-low resource medical NLP, paving the way for more inclusive, scalable, and accurate Arabic healthcare chatbot systems.

医疗AI合成数据阿拉伯语NLP大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。