arXiv:2606.13211cs.AI2026-06

系统分析医学影像AI幻觉问题,提出跨模态分类与监管兼容的应对框架。

Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints

论文配图:Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints
图 1 · 摘自论文原文
  • 构建跨模态幻觉分类体系,整合多框架覆盖成像全流程
  • 发现通用模型比医疗专用模型更少幻觉,领域微调可能引发过拟合幻觉
  • 推荐结合物理约束、思维链提示与人工审核,适配FDA全生命周期管理

AI在医学影像中的部署速度已超过对其失效模式的理解。当前临床最关切的失效形式是幻觉:生成看似合理但事实错误的结果,如虚构解剖结构、遗漏病灶、错误侧别判断及编造测量值,直接影响活检、分期和治疗方案。本文综合五年内五类影像模态的同行评审研究、基准数据集与FDA监管指引,系统分析幻觉的分类、成因、检测与缓解策略。研究回答三个问题:(1) 如何统一跨模态的幻觉分类?(2) 医疗专用基础模型为何不如通用模型抗幻觉?(3) 哪些缓解策略有效且符合FDA全生命周期监管要求?结果表明,三种分类框架联合可完整覆盖影像处理流程;通用模型在幻觉基准测试中表现更优,提示领域微调可能引入过拟合性幻觉。尽管如此,放射科医生监督仍至关重要——绝大多数AI生成警报需专家修正方可临床使用。物理启发的架构约束、思维链提示与人机协同防护各自针对不同故障模式,组合使用效果最佳。所有发现均映射至FDA的全产品生命周期(Total Product Lifecycle)与预设变更控制计划(Predetermined Change Control Plan)框架,将幻觉管理定位为持续监管义务而非仅部署前检查项。

原文摘要 · Abstract (English)

AI systems are being deployed across medical imaging faster than their failure modes are understood. At this point in time, the failure of greatest clinical concern is hallucination: clinically plausible but factually incorrect outputs, including fabricated anatomical structures, missed findings, incorrect laterality, and invented measurements in generated reports, with direct consequences, for example, for biopsy decisions, staging, and treatment planning. This structured narrative synthesizes peer-reviewed studies, benchmark datasets, and FDA regulatory guidance across five imaging modalities to produce a cross-modality analysis of hallucination taxonomy, etiology, detection, and mitigation. Specifically, we address three questions in this study: (1) how can existing taxonomies be unified across modalities?, (2) how do medical-specialized foundation models hallucinate less than general-purpose ones?, and (3) which mitigation strategies are effective and compatible with FDA lifecycle oversight? We note that three taxonomic frameworks together cover the imaging pipeline in a way no single framework does alone. We also highlight that general-purpose foundation models outperform medical-specialized models on hallucination-specific benchmarks, indicating that narrow domain fine-tuning can introduce overfitting-induced confabulation. At the same time, the oversight of radiologists remains essential; for instance, a very high percentage of of AI-generated flags required expert correction before clinical use. Physics-informed architectural constraints, Chain-of-Thought prompting, and human-in-the-loop safeguards each address different failure modes and is effective when combined. All findings are mapped to the FDA's Total Product Lifecycle and Predetermined Change Control Plan frameworks, which treat hallucination management as a lifecycle obligation rather than a pre-deployment checklist.

医学AI幻觉检测FDA监管影像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。