LLM法官难发现临床记录遗漏,新方法可精准识别。
LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It
- 重构任务:逐项核对转录内容与笔记,提升检出率
- 单次调用方法检测率36.9%,误报率仅6.2%
- 适合医疗AI质量审计,尤其需高精度遗漏检测场景
环境式AI记录员生成临床笔记,现有审计发现其主要错误为遗漏:未记录已确认的信息。当前标准由大模型法官判断,即对比笔记与录音转写并标记问题。我们质疑该法官是否能有效识别遗漏。公开语料库无法提供答案:医生参考笔记与转写存在显著差异。本研究构建500对单错误笔记,其中298个明确缺失特定事实,202个为新增或修改的对照组。在八种法官设计中,对新增/修改内容的判别率为0.79–0.94,而对遗漏的判别率仅为0.50–0.63(接近随机)。单独评估笔记时,无设计能可靠区分遗漏与完整笔记。改写、投票机制与GEPA提示优化虽调整了判别点,但未能实现可用检测。重构任务后恢复检测能力:先列出转录中确立的事实,再逐一核对笔记。两种独立方法达成目标:基于事实的流水线和经GEPA演化的一次调用提示。流水线准确命名缺失事实及其严重性,误报率2.7%;单次调用方法检测率更高(36.9% vs 24.6%,p=0.002),误报率6.2%,成本仅为十分之一。一位医师作者验证70项,当两方法分歧时,10项均支持流水线(p=0.002)。第二位非作者医师盲评严重性评分,结果一致至一个等级内。在真实厂商笔记中,基准阈值不适用,但重新校准后的单次调用方法优于八种设计,且误报率减半。若遗漏事实在其他处重述,则两种方法均失效。数据集、提示与判别结果均已开源。
原文摘要 · Abstract (English)
Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the note fails to record. The standard check is an LLM judge: a second model reads the note against the transcript and flags problems. We ask whether judges detect omissions. Public corpora cannot supply the answer key: their clinician reference notes and transcripts are materially discrepant. Our benchmark has 500 single-error note pairs from audited fact sheets, 298 with a named fact certainly absent and 202 added-or-altered controls. Across eight judge designs, paired discrimination (the flawed note below its clean twin, 0.5 a coin flip) reads 0.79-0.94 on added or altered content and 0.50-0.63 on omissions. On single notes, no design flags omissions reliably more often than perfect notes. Wording changes, voting and GEPA prompt optimisation move the operating point without creating usable detection. Restructuring the task recovers it: list the facts the transcript establishes, then check the note for each. Two methods reach it independently and trade off: a per-fact pipeline, and a GEPA-evolved prompt doing the same in one call. The pipeline's flags name the missing fact and its severity at 2.7% false alarms. The single call detects more (36.9% against 24.6%, p=0.002) at 6.2% false alarms and a tenth of the cost per note. A physician author validated 70 items and, where the two routes disagree, sided with the pipeline on 10 of 10 (p=0.002). A second clinician, not an author, graded the severity rubric blind and agrees to within a grade. On real vendor notes from a companion census no benchmark threshold transfers, but the re-calibrated single call detects more than the best of the eight at half its false-alarm rate. Omissions whose fact is restated elsewhere defeat both routes. We release the benchmark, prompts and judgements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。