评测大模型在出院后临床任务提取中的表现,发现其推理能力受限于标注规范。
Systematic Evaluation of Large Language Models for Post-Discharge Clinical Action Extraction

- 分两阶段提示框架将出院记录拆解为具体可操作任务
- 零样本大模型在二分类任务上表现接近甚至超过有监督模型
- 需标注推理过程以区分模型错误与标注差异,推动临床理解评估
本文针对出院后患者安全关键的临床任务提取问题,基于CLIP出院记录数据集,系统评估了零样本与少样本大语言模型(LLMs)的表现。为应对临床文档复杂性,提出一种两阶段抽取框架,通过分步提示策略将叙事性出院记录分解为细粒度、明确可执行的临床任务。研究贡献包括:对生成式大模型在临床任务提取中的系统评估;通用大模型与任务特定微调的BERT基模型之间的详细对比;以及不同动作类别间标注不一致性的分析。结果显示,当代大模型在二元行动力检测任务上性能可媲美或超越有监督模型,而后者在细粒度多标签分类任务中仍具显著优势,即使未进行任务特化微调且受严格数据隐私约束。定性错误分析表明,多数失败源于模型推理与数据标注规范间的错位,尤其在隐含临床行为和严格结构化标注规则情境下。这些结果表明,当前性能反映的是模型缺乏临床推理能力,而非标注本身的问题。无理由标注无法区分是模型推理失误还是标注约定差异。推进临床NLP需建立带有推理注释的数据集,记录为何某段文本被视为可操作,而不仅是标记哪些片段被标注,从而实现对模型临床理解的真实评估。
原文摘要 · Abstract (English)
The work in this paper evaluates zero-shot and few-shot large language models (LLMs) for safety-critical clinical action extraction using the CLIP discharge-note dataset, with particular emphasis on transitions of care and post-discharge patient safety. To manage the complexity of clinical documentation, we introduce a two-stage extraction framework that decomposes discharge notes, that are written in narrative form, into fine-grained, explicitly actionable clinical tasks through a staged prompting strategy. Our contributions include a systematic assessment of generative LLMs for clinical action extraction, a detailed comparison between general-purpose LLMs and task-specific supervised BERT-based models, and an analysis of annotation inconsistencies across different action categories. We show that contemporary LLMs achieve performance comparable to or exceeding supervised models on binary actionability detection, while supervised baselines retain a meaningful advantage on fine-grained multi-label category classification, despite the absence of task-specific fine-tuning and under strict data-privacy constraints. Qualitative error analysis reveals that many failures stem from misalignment between model reasoning and dataset annotation conventions, particularly in cases involving implicit clinical actions and rigid structural labeling rules. These results indicate that reported performance reflects model limitations due to lack of clinical reasoning, that is not captured by plain annotations. Labels without rationales make it impossible to distinguish clinical reasoning failures from annotation convention mismatches. Advancing clinical NLP requires reasoning-annotated datasets that document why specific spans are actionable, not merely which spans were labeled, enabling proper evaluation of model clinical understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。