arXiv:2510.12813cs.CLcs.AI2025-10被引 4

用大模型分析病历文本,自动分类癌症诊断,效果接近专家水平。

Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model Performance Evaluation Study

  • 对比4个大模型与BioBERT,用病历中的编码和自由文本分类癌症类型
  • GPT-4o在自由文本分类上表现最好,准确率达81.9%,比BioBERT高0.3个百分点
  • 模型对转移瘤和中枢神经系统肿瘤易混淆,临床应用需人工复核

电子健康记录包含结构不一或自由文本数据,需高效预处理以支持预测模型。尽管基于人工智能的自然语言处理工具在自动化诊断分类方面展现潜力,但其性能对比与临床可靠性仍需系统评估。本研究评估了4个大型语言模型(GPT-3.5、GPT-4o、Llama 3.2、Gemini 1.5)和BioBERT在从结构化与非结构化电子健康记录中分类癌症诊断的表现。分析了来自3456份癌症患者记录的762个独特诊断(326个国际疾病分类代码描述,436条自由文本条目),要求模型将诊断归入14个预设类别。两名肿瘤科专家验证分类结果。BioBERT在ICD代码上的加权宏平均F1得分为84.2,与GPT-4o的准确率90.8持平;在自由文本分类中,GPT-4o的加权宏平均F1为71.8,高于BioBERT的61.5,准确率81.9略高于81.6。GPT-3.5、Gemini和Llama整体表现较低。常见误分类包括转移瘤与中枢神经系统肿瘤混淆,以及含糊或重叠术语导致的错误。当前性能已适用于行政与研究用途,但可靠临床应用仍需标准化文档与严格人工监督。

原文摘要 · Abstract (English)

Electronic health records contain inconsistently structured or free-text data, requiring efficient preprocessing to enable predictive health care models. Although artificial intelligence-driven natural language processing tools show promise for automating diagnosis classification, their comparative performance and clinical reliability require systematic evaluation. The aim of this study is to evaluate the performance of 4 large language models (GPT-3.5, GPT-4o, Llama 3.2, and Gemini 1.5) and BioBERT in classifying cancer diagnoses from structured and unstructured electronic health records data. We analyzed 762 unique diagnoses (326 International Classification of Diseases (ICD) code descriptions, 436free-text entries) from 3456 records of patients with cancer. Models were tested on their ability to categorize diagnoses into 14predefined categories. Two oncology experts validated classifications. BioBERT achieved the highest weighted macro F1-score for ICD codes (84.2) and matched GPT-4o in ICD code accuracy (90.8). For free-text diagnoses, GPT-4o outperformed BioBERT in weighted macro F1-score (71.8 vs 61.5) and achieved slightly higher accuracy (81.9 vs 81.6). GPT-3.5, Gemini, and Llama showed lower overall performance on both formats. Common misclassification patterns included confusion between metastasis and central nervous system tumors, as well as errors involving ambiguous or overlapping clinical terminology. Although current performance levels appear sufficient for administrative and research use, reliable clinical applications will require standardized documentation practices alongside robust human oversight for high-stakes decision-making.

癌症诊断大模型病历分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。