医学影像模型常胡说八道,这篇论文系统评测并提出防范方案。
HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation

- 构建了856张胸片的基准测试,评估六种模型在三种问题下的表现
- 发现61.9%至82.3%输出含幻觉,严重错误达80.2%
- 通过模型集成和回答长度监控,可大幅降低幻觉风险
视觉语言模型(VLMs)在医学影像解读中日益普及,但常产生临床看似合理却事实错误的描述,直接威胁患者安全。本文提出HalluCXR基准,评估六种架构各异的VLMs在856张分层的MIMIC-CXR胸片上的表现,覆盖三种查询类型,共生成15,408次模型输出。建立八类幻觉分类体系并附临床严重度评分,通过两层检测流程验证,对比250条人工标注(自动检测F1=0.959;LLM判别F1=0.907)。结果显示,61.9%–82.3%的输出含幻觉,最严重错误占比达80.2%。关键模式包括:正常胸片反而最易引发严重幻觉,常见病灶被系统性夸大而罕见病灶常被遗漏,且回答长度单独可预测幻觉风险(AUC最高达0.908)。六模型集成可减少84.8%的虚假生成,代价是漏诊增多;三模型子集则在成本减半情况下保持相近性能。结果表明,幻觉审计、基于冗长度的风险监测与集成式安全层是临床部署的必要前提。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly used for medical image interpretation, yet they frequently hallucinate, generating clinically plausible but factually incorrect findings that pose direct patient safety risks. We introduce HalluCXR, a benchmark evaluating six architecturally diverse VLMs across 856 stratified MIMIC-CXR chest radiographs and three query types, yielding 15,408 model evaluations. An eight-category hallucination taxonomy with clinical severity ratings and a two-layer detection pipeline are validated against 250 human annotations (auto-detection F1=0.959; LLM judge F1=0.907). We find that 61.9--82.3% of outputs contain hallucinations, with clinically dangerous errors in up to 80.2%. Three key patterns emerge: normal radiographs paradoxically attract the most severe hallucinations, common findings are systematically over-fabricated while rare findings go under-detected, and response length alone predicts hallucination risk (AUC up to 0.908). A six-model ensemble reduces fabrication by up to 84.8% at the cost of increased omission; a three-model subset retains comparable performance at half the cost. These results establish that hallucination auditing, verbosity-based risk monitoring, and ensemble-based safety layers are prerequisites for clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。