CHECK让医学大模型持续检测并消除幻觉,显著提升可信度。
Trustworthy AI for Medicine: Continuous Hallucination Detection and Elimination with CHECK
- 用信息论构建分类器,结合临床数据库实时识别事实与推理幻觉。
- 将幻觉率从31%降至0.3%,在多个医学测试中达到92.1%的通过率。
- 适合医疗AI部署者、研究人员及关注高风险领域可信生成的人群。
大型语言模型在医疗领域前景广阔,但幻觉仍是临床应用的主要障碍。本文提出CHECK框架,通过融合结构化临床数据库与基于信息论的分类器,实现对事实性与推理性幻觉的持续检测。在100项关键临床试验的1500个问题上评估,该方法将LLama3.3-70B-Instruct的幻觉率从31%降至0.3%,使开源模型达到业界领先水平。其分类器在多个医学基准上表现优异,AUC达0.95–0.96,包括MedQA(USMLE)和HealthBench真实多轮医疗问答。利用幻觉概率指导GPT-4o优化,并合理调度计算资源,使其在USMLE考试中通过率提升5个百分点,达到92.1%的最新纪录。通过将幻觉抑制至临床可接受误差阈值以下,CHECK为医疗及其他高风险领域的大模型安全部署提供了可扩展基础。
原文摘要 · Abstract (English)
Large language models (LLMs) show promise in healthcare, but hallucinations remain a major barrier to clinical use. We present CHECK, a continuous-learning framework that integrates structured clinical databases with a classifier grounded in information theory to detect both factual and reasoning-based hallucinations. Evaluated on 1500 questions from 100 pivotal clinical trials, CHECK reduced LLama3.3-70B-Instruct hallucination rates from 31% to 0.3% - making an open source model state of the art. Its classifier generalized across medical benchmarks, achieving AUCs of 0.95-0.96, including on the MedQA (USMLE) benchmark and HealthBench realistic multi-turn medical questioning. By leveraging hallucination probabilities to guide GPT-4o's refinement and judiciously escalate compute, CHECK boosted its USMLE passing rate by 5 percentage points, achieving a state-of-the-art 92.1%. By suppressing hallucinations below accepted clinical error thresholds, CHECK offers a scalable foundation for safe LLM deployment in medicine and other high-stakes domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。