用多层大模型框架提升阿拉伯语医疗文本疾病预测准确率
An Ensemble Classification Approach in A Multi-Layered Large Language Model Framework for Disease Prediction
- 用三种预处理方法+三个阿拉伯语模型,再用投票融合提升鲁棒性
- 最高准确率达80.56%,验证了多源文本与模型融合的有效性
- 首个结合大模型预处理与阿拉伯语模型的疾病预测框架,适合医疗文本分析者
社交远程医疗在医疗领域取得显著进展,使患者能在线发布症状并参与诊疗。用户常在社交媒体和健康平台发布症状,形成庞大的医疗数据资源,可用于疾病分类。大型语言模型(如LLAMA3、GPT-3.5)及基于Transformer的模型(如BERT)在处理复杂医疗文本方面表现出色。本研究评估了三种阿拉伯语医疗文本预处理方法:摘要、精炼和命名实体识别(NER),并在微调后的阿拉伯语Transformer模型(CAMeLBERT、AraBERT、AsafayaBERT)上应用。为增强鲁棒性,采用多数投票集成方法,融合原始文本与预处理后文本表示的预测结果。该方法达到最高分类准确率80.56%,证明了利用多种文本表示与模型预测可有效提升对医疗文本的理解。据我们所知,这是首个将基于大模型的预处理与微调的阿拉伯语Transformer模型及集成学习相结合,用于阿拉伯语社交远程医疗数据中的疾病分类工作。
原文摘要 · Abstract (English)
Social telehealth has made remarkable progress in healthcare by allowing patients to post symptoms and participate in medical consultations remotely. Users frequently post symptoms on social media and online health platforms, creating a huge repository of medical data that can be leveraged for disease classification. Large language models (LLMs) such as LLAMA3 and GPT-3.5, along with transformer-based models like BERT, have demonstrated strong capabilities in processing complex medical text. In this study, we evaluate three Arabic medical text preprocessing methods such as summarization, refinement, and Named Entity Recognition (NER) before applying fine-tuned Arabic transformer models (CAMeLBERT, AraBERT, and AsafayaBERT). To enhance robustness, we adopt a majority voting ensemble that combines predictions from original and preprocessed text representations. This approach achieved the best classification accuracy of 80.56%, thus showing its effectiveness in leveraging various text representations and model predictions to improve the understanding of medical texts. To the best of our knowledge, this is the first work that integrates LLM-based preprocessing with fine-tuned Arabic transformer models and ensemble learning for disease classification in Arabic social telehealth data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。