arXiv:2502.00063cs.CLcs.AI2025-02被引 6

用多层大模型框架提升阿拉伯语医疗文本疾病预测准确率

A Multi-Layered Large Language Model Framework for Disease Prediction

  • 融合文本摘要、优化与命名实体识别的三层预处理流程
  • 采用CAMeL-BERT+NER增强文本,疾病类型分类率达83%
  • 适合关注低资源语言医疗AI的开发者与研究者

社交远程医疗通过让用户在社交媒体和在线健康平台分享症状,彻底改变了医疗模式,形成了海量可利用的医疗数据。大型语言模型(如LLAMA3、GPT-3.5 Turbo、BERT)能处理复杂医学文本,提升疾病分类能力。本研究探索了三种阿拉伯语医学文本预处理方法:文本摘要、文本优化和命名实体识别(NER)。对比评估了CAMeL-BERT、AraBERT和Asafaya-BERT配合LoRA微调的效果,结果显示使用经过NER增强的文本时,CAMeL-BERT表现最佳,疾病类型分类准确率为83%,严重程度评估达69%。未微调模型表现较差(类型分类13%-20%,严重程度40%-49%)。将大模型融入社交远程医疗系统可显著提升诊断准确性和治疗效果。

原文摘要 · Abstract (English)

Social telehealth has revolutionized healthcare by enabling patients to share symptoms and receive medical consultations remotely. Users frequently post symptoms on social media and online health platforms, generating a vast repository of medical data that can be leveraged for disease classification and symptom severity assessment. Large language models (LLMs), such as LLAMA3, GPT-3.5 Turbo, and BERT, process complex medical data to enhance disease classification. This study explores three Arabic medical text preprocessing techniques: text summarization, text refinement, and Named Entity Recognition (NER). Evaluating CAMeL-BERT, AraBERT, and Asafaya-BERT with LoRA, the best performance was achieved using CAMeL-BERT with NER-augmented text (83% type classification, 69% severity assessment). Non-fine-tuned models performed poorly (13%-20% type classification, 40%-49% severity assessment). Integrating LLMs into social telehealth systems enhances diagnostic accuracy and treatment outcomes.

疾病预测大模型阿拉伯语医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。