arXiv:2502.07755cs.CLcs.AI2025-02被引 4

用DeBERTa+动态位置门控提升医疗文本诊断准确率

An Advanced NLP Framework for Automated Medical Diagnosis with DeBERTa and Dynamic Contextual Positional Gating

  • 结合回译增强数据,缓解过拟合问题
  • 在医疗文本分类中达99.78%准确率
  • 适合需要高精度的临床辅助诊断场景

本文提出一种新型自然语言处理框架,用于提升医疗诊断的自动化水平。通过回译技术生成多样化改写数据集,增强分类任务的鲁棒性并缓解过拟合。采用引入解码增强与解耦注意力的DeBERTa模型,并结合动态上下文位置门控(DCPG),根据语义上下文动态调整位置信息影响,生成高质量文本嵌入。分类阶段使用基于注意力的前馈神经网络(ABFNN),聚焦关键特征以提升决策准确性。该架构应用于症状、临床记录等医疗文本分类,有效应对医疗数据复杂性。实验表明,该方法在多项指标上表现优异:准确率99.78%、召回率99.72%、精确率99.79%、F1分数99.75%,显著优于现有方法,展现出在自动化诊断与临床决策支持中的巨大潜力。

原文摘要 · Abstract (English)

This paper presents a novel Natural Language Processing (NLP) framework for enhancing medical diagnosis through the integration of advanced techniques in data augmentation, feature extraction, and classification. The proposed approach employs back-translation to generate diverse paraphrased datasets, improving robustness and mitigating overfitting in classification tasks. Leveraging Decoding-enhanced BERT with Disentangled Attention (DeBERTa) with Dynamic Contextual Positional Gating (DCPG), the model captures fine-grained contextual and positional relationships, dynamically adjusting the influence of positional information based on semantic context to produce high-quality text embeddings. For classification, an Attention-Based Feedforward Neural Network (ABFNN) is utilized, effectively focusing on the most relevant features to improve decision-making accuracy. Applied to the classification of symptoms, clinical notes, and other medical texts, this architecture demonstrates its ability to address the complexities of medical data. The combination of data augmentation, contextual embedding generation, and advanced classification mechanisms offers a robust and accurate diagnostic tool, with potential applications in automated medical diagnosis and clinical decision support. This method demonstrates the effectiveness of the proposed NLP framework for medical diagnosis, achieving remarkable results with an accuracy of 99.78%, recall of 99.72%, precision of 99.79%, and an F1-score of 99.75%. These metrics not only underscore the model's robust performance in classifying medical texts with exceptional precision and reliability but also highlight its superiority over existing methods, making it a highly promising tool for automated diagnostic systems.

医疗NLPDeBERTa精准诊断文本分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。