大规模评估大模型重写临床笔记的质量,发现其保留核心信息但丢失细节。
Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale
- 分块重写可减少信息丢失,但可能影响事实准确性。
- 生成笔记保留粗粒度任务预测能力,但影响精细编码如ICD。
- 适合用于罕见疾病编码的数据增强,尤其在标注数据稀缺时。
大型语言模型(LLMs)可用于生成或合成临床文本,应用于临床文档改进与文本分析增补。然而,现有评估多聚焦单一维度(如相似性或实用性),而这些维度应并行考察。本研究在百万级MIMIC数据库临床笔记基础上,系统评估了大模型生成文本的内在质量、外在效用及事实正确性。结果表明,尽管语言表达发生显著变化,合成笔记仍能保持核心临床信息和粗粒度任务的预测能力,但在精细任务(如ICD编码)中丢失细节。通过分块重写可有效缓解该问题,但因上下文不完整导致事实精度下降。错误分析显示,主要错误类型为临床语境误读、时间混淆、测量偏差及虚构陈述。最后,即使合成数据无任务特定设计,仍可有效增强罕见ICD代码的任务训练。
原文摘要 · Abstract (English)
Large language models (LLMs) can generate or synthesize clinical text for a wide range of applications, from improving clinical documentation to augmenting clinical text analytics. Yet evaluations typically focus on a narrow aspect -- such as similarity or utility comparisons -- even though these aspects are complementary and best viewed in parallel. In this study, we aim to conduct a systematic evaluation of LLM-generated clinical text, which includes intrinsic, extrinsic, and factuality evaluations of synthetic clinical notes rephrased from MIMIC databases at million-note scale. Our analysis demonstrates that synthetic notes preserve core clinical information and predictive utility for coarse-grained tasks despite substantial linguistic changes, but lose fine-grained details for task like ICD coding. We show this loss of detail can be substantially mitigated by rephrasing notes by chunks rather than by the whole note, but at the cost of reduced factual precision under incomplete context. Through fact-checking and error analysis, we further find that synthesis errors are dominated by misinterpretation of clinical context, alongside temporal confusion, measurement errors, and fabricated claims. Finally, we show that the synthetic notes -- despite their task-agnostic nature -- can effectively augment task-specific training for rare ICD codes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。