arXiv:2409.20147cs.CLcs.AI2024-09被引 2

在小样本非英语医学文本分类中,领域预训练BERT表现最佳。

Classification of Radiological Text in Small and Imbalanced Datasets in a Non-English Language

  • 使用领域内预训练的BERT模型处理丹麦语放射科报告
  • 模型准确率最高达78.3%,但全无监督仍不达标
  • 适合低资源语言医疗文本的初步筛选与降本标注

医学自然语言处理在小规模、非英语、标注样本少且类别不平衡的数据集上表现不佳,当前尚无统一解决方案。我们评估了多种NLP模型,包括BERT类变换器、基于句子嵌入的少样本学习(SetFit)以及提示式大语言模型(LLM),在三种丹麦语癫痫患者磁共振影像报告数据集上的表现。结果表明,针对放射科报告领域预训练的BERT类模型在此场景下表现最优。值得注意的是,SetFit和LLM模型表现均逊于BERT类模型,其中LLM表现最差。重要的是,所研究的模型均无法在无监督条件下达到足够准确率以直接用于文本分类。然而,它们在数据过滤方面展现出潜力,可显著减少人工标注所需工作量。

原文摘要 · Abstract (English)

Natural language processing (NLP) in the medical domain can underperform in real-world applications involving small datasets in a non-English language with few labeled samples and imbalanced classes. There is yet no consensus on how to approach this problem. We evaluated a set of NLP models including BERT-like transformers, few-shot learning with sentence transformers (SetFit), and prompted large language models (LLM), using three datasets of radiology reports on magnetic resonance images of epilepsy patients in Danish, a low-resource language. Our results indicate that BERT-like models pretrained in the target domain of radiology reports currently offer the optimal performances for this scenario. Notably, the SetFit and LLM models underperformed compared to BERT-like models, with LLM performing the worst. Importantly, none of the models investigated was sufficiently accurate to allow for text classification without any supervision. However, they show potential for data filtering, which could reduce the amount of manual labeling required.

医学NLP小样本学习低资源语言文本分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。