AI重写放射科报告看似更规范,实则可能让文本与影像脱节,反而影响医疗AI训练效果。
The Slop Paradox: How Synthetic Standardization Erodes Clinical Uncertainty and Cross-Modal Alignment in AI-Rewritten Radiology Reports
- 用三种真实场景的LLM重写任务测试报告信息损失
- 标准化重写虽保留更多医学实体,但使图文对齐下降14.9%-16.5%
- 罕见病未被更严重破坏,提示问题难通过特定监测发现
AI辅助临床文书工具越来越多地使用大语言模型(LLMs)对放射科报告进行摘要、标准化和重写。我们基于印第安纳大学450份胸部X光报告,通过三种真实场景的LLM重写任务——电子病历摘要、标准格式重写和教学案例准备——开展受控测量。评估指标包括医学实体侵蚀(通过医学命名实体识别)、不确定性表达衰减(语义模糊性减少)以及跨模态对齐退化(通过BiomedCLIP图像-文本相似度)。核心发现是:信息损失与跨模态保真度之间存在分离。电子病历摘要在内容层面破坏最严重,导致51.4%的临床实体丢失、43.7%的不确定性语言消失,但图像-文本对齐仅下降2.5%;而旨在生成更干净训练数据的标准重写和教学案例准备,虽仅侵蚀26.8%和29.3%的实体,却造成14.9%-16.5%的对齐下降,是前者的六到七倍。这一矛盾现象称为‘废料悖论’:让文本看起来更适合多模态训练的重写方式,恰恰使其远离原始图像。与预设假设相反,罕见病并未被更严重侵蚀——九组常见病与罕见病对比中,无一显著差异,且名义差异反向显示常见病受影响更大,说明污染难以通过病种特异性监控发现。决定降级的主要因素是重写任务类型,而非临床内容本身。该研究对多模态医疗AI数据集构建及AI辅助文书治理具有重要意义。
原文摘要 · Abstract (English)
AI-assisted clinical documentation tools increasingly summarize, standardize, and reformat radiology reports using large language models (LLMs). We present a controlled measurement of the resulting information degradation. Using 450 chest X-ray reports from the Indiana University dataset, we generate synthetic versions via three realistic LLM rewriting tasks: EHR summarization, standardized rewriting, and teaching case preparation. We measure entity erosion (via medical NER), hedging collapse (loss of clinical uncertainty language), and cross-modal alignment degradation (via BiomedCLIP image-text similarity). Our central finding is a dissociation between information loss and cross-modal fidelity. EHR summarization is the most destructive at the content level, eroding 51.4% of clinical entities and 43.7% of hedging language, yet it preserves image-text alignment almost entirely (a 2.5% drop). The two tasks meant to produce cleaner training data, standardized rewriting and teaching case preparation, do the reverse: they preserve more entities (26.8% and 29.3% eroded) but cause 14.9-16.5% alignment drops, six to seven times those of EHR summarization. We term this the slop paradox: rewriting that makes clinical text look cleaner for multimodal training is precisely what pulls it away from the image. Contrary to our pre-specified hypothesis, rare pathologies were not preferentially degraded: across nine rare-versus-common comparisons, no difference survived multiple-comparison correction, and nominal differences ran in the opposite direction (common > rare), so contamination is invisible to condition-specific monitoring. The dominant determinant of degradation is the type of AI rewriting task, not the clinical content. These findings bear on multimodal medical AI dataset construction and the governance of AI-assisted clinical documentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。