arXiv:2508.07031eess.IVcs.AI2025-08ICCV被引 13

研究大模型在医学影像中生成错误报告与图像的幻觉问题

Trustworthy Medical Imaging with Large Language Models: A Study of Hallucinations Across Modalities

  • 分析大模型从影像生成报告和从文本生成影像的幻觉机制
  • 发现两种任务中均存在事实矛盾与解剖错误,影响临床可靠性
  • 适合关注AI医疗安全与可信度的研究者与临床开发者

大型语言模型(LLMs)正被广泛应用于医学影像解读与合成图像生成。然而,这些模型常产生幻觉——即自信但错误的输出,可能误导临床决策。本研究考察了两个方向的幻觉:影像到文本(从X光、CT或MRI生成报告),以及文本到影像(根据临床提示生成医学图像)。通过专家定义的标准评估不同模态下的错误,如事实不一致与解剖错误。研究发现,在解释与生成任务中均存在普遍的幻觉模式,对临床可靠性构成威胁。同时探讨了模型架构与训练数据等因素对失败的影响。系统性研究双向任务,为提升基于大模型的医学影像系统安全性与可信度提供洞见。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly applied to medical imaging tasks, including image interpretation and synthetic image generation. However, these models often produce hallucinations, which are confident but incorrect outputs that can mislead clinical decisions. This study examines hallucinations in two directions: image to text, where LLMs generate reports from X-ray, CT, or MRI scans, and text to image, where models create medical images from clinical prompts. We analyze errors such as factual inconsistencies and anatomical inaccuracies, evaluating outputs using expert informed criteria across imaging modalities. Our findings reveal common patterns of hallucination in both interpretive and generative tasks, with implications for clinical reliability. We also discuss factors contributing to these failures, including model architecture and training data. By systematically studying both image understanding and generation, this work provides insights into improving the safety and trustworthiness of LLM driven medical imaging systems.

医学影像大模型幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。