医学病历生成评估需重定义幻觉,避免误判有效临床推理。
Beyond Literal Summarization: Redefining Hallucination for Medical SOAP Note Evaluation

- 用医学本体检索与校准提示,让评估更符合临床思维。
- 传统方法误判率35%,新方法降至9%。
- 适合医疗AI评估、临床语言模型开发者参考。
评估大语言模型在临床文档任务(如SOAP病历生成)中的表现仍具挑战性。与标准摘要不同,此类任务要求临床抽象、口语化表达的规范化及基于医学知识的推理。然而,当前评估方法(如自动指标和以LLM为裁判的框架)依赖词汇忠实度,常将未在原文明确出现的信息标记为幻觉。我们发现,这类方法会系统性地将合法的临床输出误判为错误,夸大幻觉率并扭曲模型评估。分析表明,许多被标记为幻觉的内容实为有效的临床转换,包括同义词映射、检查结果抽象、诊断推断及符合指南的治疗规划。通过结合医学本体的检索与校准提示,使评估标准契合临床推理,结果显示:在词汇评估下,平均幻觉率为35%,严重惩罚有效推理;而在考虑推理的评估下,该值降至9%,剩余案例反映真正的安全问题。研究揭示,现有评估方法过度惩罚合理临床推理,可能衡量的是评估设计的偏差而非真实错误,强调在高上下文领域如医学中需要临床导向的评估范式。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) for clinical documentation tasks such as SOAP note generation remains challenging. Unlike standard summarization, these tasks require clinical abstraction, normalization of colloquial language, and medically grounded inference. However, prevailing evaluation methods including automated metrics and LLM as judge frameworks rely on lexical faithfulness, often labeling any information not explicitly present in the transcript as hallucination. We show that such approaches systematically misclassify clinically valid outputs as errors, inflating hallucination rates and distorting model assessment. Our analysis reveals that many flagged hallucinations correspond to legitimate clinical transformations, including synonym mapping, abstraction of examination findings, diagnostic inference, and guideline consistent care planning. By aligning evaluation criteria with clinical reasoning through calibrated prompting and retrieval grounded in medical ontologies we observe a significant shift in outcomes. Under a lexical evaluation regime, the mean hallucination rate is 35%, heavily penalizing valid reasoning. With inference aware evaluation, this drops to 9%, with remaining cases reflecting genuine safety concerns. These findings suggest that current evaluation practices over penalize valid clinical reasoning and may measure artifacts of evaluation design rather than true errors, underscoring the need for clinically informed evaluation in high context domains like medicine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。