研究大模型在病历提取中对提示、模型和模板的敏感性,发现模型影响远大于提示。
Measuring the sensitivity of LLM-based structured extraction to prompt, model, and schema choices in clinical discharge summaries

- 固定任务与数据,逐一测试提示、模型、模板变化对输出的影响。
- 模型大小改变导致近半病历分类标签重分配,提示变化影响较小。
- 分歧主要来自‘未记录’与‘无’的区分,而非是否存在临床特征。
大型语言模型越来越多用于从临床自由文本中提取结构化信息,但其输出对上游配置选择的敏感性尚不明确。本研究在无人工标注真值的情况下,通过固定提取任务并逐项调整提示、模型或模板,测量其敏感性。采用包含17个三分类(是/否/未记录)标志和47类主因标签的固定模板,在MIMIC-IV v3.1的出院摘要上,分别用两种模型尺寸运行三种提示变体。基于ICD分层子集计算提示间一致性(Cohen's kappa)。通过同一病历的配对比较分离模型影响,并将三分类退化为二分类以检验模板贡献。三分类标志上,两模型的总体交叉提示一致性相近(中位kappa分别为0.69和0.68),大模型在部分字段提升一致性,其他则下降,体现的是再分配而非无影响。将三分类退化为二分类后,大部分提示分歧消失,问题集中于“不存在”与“未记录”的区分。在多分类主因分类中,更换模型使近半数病历的主导标签发生变化,而提示变化仅影响约八分之一;大模型显著减少对残余归档类别(catch-all)的依赖(从44%降至26%)。结果表明,分歧主要源于模板设计中的“不存在”与“未记录”边界,且模型选择对多分类影响远大于提示表达。该方法可重复用于大规模部署中的提取一致性审计。
原文摘要 · Abstract (English)
Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream configuration choices is less understood than their accuracy on fixed benchmarks. This work measures that sensitivity without human-annotated ground truth, by holding the extraction task fixed and varying one choice at a time. The fixed schema comprises 17 clinical documentation flags on a three-way yes/no/not_documented value set and a 47-tag vocabulary for the primary admission reason. Three prompt variants expressing this schema were each run at two model sizes on MIMIC-IV v3.1 discharge summaries. Cross-prompt agreement was measured by Cohen's kappa on ICD-stratified subsets. A paired same-note comparison isolated the effect of model choice, and a post-hoc collapse of the three-way flags to binary tested the schema's contribution to disagreement. On the three-way flags, the two models reach the same pooled cross-prompt agreement (median kappa 0.69 and 0.68); the larger model raises agreement on some fields and lowers it on others, a redistribution rather than the absence of an effect. Collapsing the schema to binary dissolves most of the cross-prompt disagreement, locating it on the absence-versus-silence distinction rather than on whether the finding is present. On the multi-class admission categorization, changing the model reassigns the dominant tag on close to half of all notes while changing the prompt phrasing reassigns it on roughly one in eight, and the larger model places far less mass on residual catch-all categories (44% to 26%). These patterns indicate a schema-imposed source of disagreement concentrated on the absence-versus-silence axis and a dominance of model over prompt phrasing on multi-class categorization, identified by a reusable methodology for auditing extraction reproducibility on a population-scale deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。