审计三款商用AI病历记录员,发现每三份病历就有一份存在可验证错误。
One note in three: a verified census of three deployed AI scribes, and the instrument that counted it
- 通过多轮对抗性审查,从565份病历中筛选出618个真实错误
- 31.3%的病历存在错误,主要集中在过敏史、用药信息和虚构患者身份
- 错误率受审核标准影响极大,不同方法结果差异可达28%至97%
我们对三款商业AI病历记录员在相同142次诊疗中的表现进行了审计:涵盖565份英国初级保健与美国门诊录音及自编场景的病历。经过十二轮探索提出13,678个候选错误,经重要性筛选后5,898个进入由两个不同模型家族组成的对抗评审组,最终618个存活。每三份病历中就有一份(31.3% [27.0, 35.6])含有可验证错误,主要集中于过敏史、药物信息、虚构患者身份,以及电话问诊中将病史误写为检查内容。未使用患者档案的情况下,错误率为24.8% [20.8, 29.0]。此外,还发现一种未被现有分类涵盖的错误类型:医生撤回的治疗被记录为已实施。两名临床医生盲评独立样本,医师作者确认20/21(95.2% [77.3, 99.2]),独立临床专家确认12/12([75.8, 100]),所有拒绝均判定为真实。错误率不仅取决于记录员,更受评估工具影响:固定模型与设置下,仅改变评审指令,验证率从9.3%升至79.0%;不同模型家族间差异亦显著,温和模型验证率54.8%对比激进模型27.8%。根据标准不同,错误率在28%至97%之间浮动。现有审计结果差异巨大,仅仪器差异即可解释部分分歧:其报告的遗漏错误占比为54-86%,而本研究为23.1%。所有618个发现连同原文档证据、提示词、模型版本及可复现流水线均已公开。
原文摘要 · Abstract (English)
Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error taxonomies: a treatment the clinician retracts, recorded as delivered care. Two clinicians adjudicated blind, disjoint samples: a physician author upheld 20 of 21 findings (95.2% [77.3, 99.2]) and an independent clinician, not an author, 12 of 12 ([75.8, 100]); both judged every sampled refusal genuine. A failure rate depends on the instrument as much as the scribes. With model, evidence and settings fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the reviewing family moves it too: alone at that instruction the gentler flags 54.8% of notes against 27.8%. Between 28% and 97% of sampled notes carry a failure depending on the standard. Published audits disagree among themselves by a margin instrument differences alone can produce: omission is 54-86% of their errors against our 23.1%. We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。