生成可控幻觉图像,用于评估医疗影像修复模型的可靠性。
HalluGen: Synthesizing Realistic and Controllable Hallucinations for Evaluating Image Restoration
- 用扩散模型合成类型、位置、严重程度可控的幻觉图像。
- 构建包含4350张标注图的首个大规模幻觉数据集,提升评估精度。
- 适合研究幻觉检测、医学影像修复与模型可信度的学者使用。
生成模型容易产生幻觉:即在真实图像中不存在但看起来合理的结构。这一问题在医疗影像、工业检测和遥感等安全关键领域尤为严重,可能导致误诊。例如,在资源受限地区广泛使用的低场磁共振成像中,修复模型虽能提升图像质量,但幻觉可能引发严重诊断错误。评估幻觉一直受限于循环依赖:需标注数据,但标注成本高且主观性强。本文提出HalluGen,一种基于扩散模型的框架,可合成感知真实但语义错误的幻觉图像(分割交并比从0.86降至0.36),生成4,350张标注图像,源自1,450张脑部MRI图像,用于低场增强任务。该数据集支持系统性评估幻觉检测与缓解方法。我们将其应用于两项任务:(1) 基准测试图像质量指标,提出基于特征评估的语义幻觉评估方法(SHAFE),通过软注意力池化提升幻觉敏感性;(2) 训练无需参考的幻觉检测器,可泛化至真实修复失败场景。HalluGen与开放数据集为安全关键图像修复中的幻觉评估提供了首个可扩展基础。
原文摘要 · Abstract (English)
Generative models are prone to hallucinations: plausible but incorrect structures absent in the ground truth. This issue is problematic in image restoration for safety-critical domains such as medical imaging, industrial inspection, and remote sensing, where such errors undermine reliability and trust. For example, in low-field MRI, widely used in resource-limited settings, restoration models are essential for enhancing low-quality scans, yet hallucinations can lead to serious diagnostic errors. Progress has been hindered by a circular dependency: evaluating hallucinations requires labeled data, yet such labels are costly and subjective. We introduce HalluGen, a diffusion-based framework that synthesizes realistic hallucinations with controllable type, location, and severity, producing perceptually realistic but semantically incorrect outputs (segmentation IoU drops from 0.86 to 0.36). Using HalluGen, we construct the first large-scale hallucination dataset comprising 4,350 annotated images derived from 1,450 brain MR images for low-field enhancement, enabling systematic evaluation of hallucination detection and mitigation. We demonstrate its utility in two applications: (1) benchmarking image quality metrics and developing Semantic Hallucination Assessment via Feature Evaluation (SHAFE), a feature-based metric with soft-attention pooling that improves hallucination sensitivity over traditional metrics; and (2) training reference-free hallucination detectors that generalize to real restoration failures. Together, HalluGen and its open dataset establish the first scalable foundation for evaluating hallucinations in safety-critical image restoration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。