arXiv:2509.26136cs.CL2025-09Conference of the …被引 2

首个对比生成与编码器模型的临床诊断预测基准,揭示编码器更优。

CliniBench: A Clinical Outcome Prediction Benchmark for Generative and Encoder-Based Language Models

  • 构建MIMIC-IV数据集上诊断预测的统一评估框架。
  • 12个生成模型均弱于3个编码器模型,准确率差距显著。
  • 检索增强可提升生成模型表现,适合临床应用研究者参考。

随着能力提升,生成式大语言模型(LLMs)正被探索用于复杂医疗任务,但其在真实临床中的有效性仍待验证。为此,我们提出CliniBench,首个可在MIMIC-IV数据集上对经典编码器分类器与生成式LLMs进行出院诊断预测比较的基准。我们系统评估了12个生成式LLMs和3个编码器分类器,结果表明编码器模型在诊断预测中始终优于生成模型。我们还测试了多种基于相似患者检索的上下文学习增强策略,发现其能显著提升生成模型性能。

原文摘要 · Abstract (English)

With their growing capabilities, generative large language models (LLMs) are being increasingly investigated for complex medical tasks. However, their effectiveness in real-world clinical applications remains underexplored. To address this, we present CliniBench, the first benchmark that enables comparability of well-studied encoder-based classifiers and generative LLMs for discharge diagnosis prediction from admission notes in MIMIC-IV dataset. Our extensive study compares 12 generative LLMs and 3 encoder-based classifiers and demonstrates that encoder-based classifiers consistently outperform generative models in diagnosis prediction. We assess several retrieval augmentation strategies for in-context learning from similar patients and find that they provide notable performance improvements for generative LLMs.

临床预测大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。