arXiv:2507.03001cs.CLcs.AI2025-07被引 3

测试大模型从病历中分类疾病编码的准确率,发现仍难替代人工。

Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs

  • 用临床NLP工具提取术语,统一提示格式评估11个大模型
  • 最高F1仅57%,越具体的编码准确率越低
  • 带推理能力的模型表现更好,适合辅助医生而非完全自动化

本研究评估大语言模型(LLMs)从医院出院小结中分类ICD-10编码的能力,该任务在医疗中至关重要但易出错。基于MIMIC-IV数据集中的1,500份摘要,聚焦10个最常见ICD-10编码,测试了11个LLM,包括具备与不具备结构化推理能力的模型。使用cTAKES提取医学术语,并以统一的编码员风格提示模型。所有模型的F1分数均未超过57%,且随着编码特异性增加,性能下降。具备推理能力的模型总体表现优于非推理模型,其中Gemini 2.5 Pro表现最佳。部分编码如慢性心脏病相关编码分类更准确。结果表明,尽管LLMs可辅助人工编码,但尚不足以实现全自动应用。未来工作应探索混合方法、领域特定模型训练及结构化临床数据的应用。

原文摘要 · Abstract (English)

This study evaluates how well large language models (LLMs) can classify ICD-10 codes from hospital discharge summaries, a critical but error-prone task in healthcare. Using 1,500 summaries from the MIMIC-IV dataset and focusing on the 10 most frequent ICD-10 codes, the study tested 11 LLMs, including models with and without structured reasoning capabilities. Medical terms were extracted using a clinical NLP tool (cTAKES), and models were prompted in a consistent, coder-like format. None of the models achieved an F1 score above 57%, with performance dropping as code specificity increased. Reasoning-based models generally outperformed non-reasoning ones, with Gemini 2.5 Pro performing best overall. Some codes, such as those related to chronic heart disease, were classified more accurately than others. The findings suggest that while LLMs can assist human coders, they are not yet reliable enough for full automation. Future work should explore hybrid methods, domain-specific model training, and the use of structured clinical data.

医疗AI大模型疾病编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。