arXiv:2506.01921cs.CVcs.AI2025-06EMNLP被引 6

首个医学影像文本编辑可靠性评测基准,解决临床应用难题

MedEBench: Diagnosing Reliability in Text-Guided Medical Image Editing

  • 构建1182对临床级图像-提示数据,覆盖70项任务与13个解剖区域
  • 发现7个主流模型普遍存在定位错误,平均注意力错位率超40%
  • 提出基于注意力对齐的诊断方法,用IoU衡量模型关注区域准确性

文本引导的图像编辑在自然图像领域进展显著,但在医学影像中应用受限且缺乏标准化评估框架。此类技术可革新临床实践,实现个性化手术规划、提升医学教育质量并改善医患沟通。为弥合这一差距,我们提出MedEBench1,一个用于诊断文本引导医学图像编辑可靠性的基准。MedEBench包含1,182对经临床审核的图像-提示对,涵盖70种不同编辑任务和13个解剖区域。其贡献包括:(1)基于临床需求的评估框架,量化编辑精度、上下文保留度与视觉质量,并提供详细的编辑意图描述及对应感兴趣区域(ROI)掩码;(2)对七种前沿模型的全面比较,揭示其一致性的失败模式;(3)一种诊断性错误分析技术,通过计算模型注意力图与ROI掩码间的交并比(IoU),识别出模型将注意力误定位至错误解剖区域的问题。MedEBench为开发更可靠、更具临床价值的文本引导医学图像编辑工具奠定基础。

原文摘要 · Abstract (English)

Text-guided image editing has seen significant progress in natural image domains, but its application in medical imaging remains limited and lacks standardized evaluation frameworks. Such editing could revolutionize clinical practices by enabling personalized surgical planning, enhancing medical education, and improving patient communication. To bridge this gap, we introduce MedEBench1, a robust benchmark designed to diagnose reliability in text-guided medical image editing. MedEBench consists of 1,182 clinically curated image-prompt pairs covering 70 distinct editing tasks and 13 anatomical regions. It contributes in three key areas: (1) a clinically grounded evaluation framework that measures Editing Accuracy, Context Preservation, and Visual Quality, complemented by detailed descriptions of intended edits and corresponding Region-of-Interest (ROI) masks; (2) a comprehensive comparison of seven state-of-theart models, revealing consistent patterns of failure; and (3) a diagnostic error analysis technique that leverages attention alignment, using Intersection-over-Union (IoU) between model attention maps and ROI masks to identify mislocalization issues, where models erroneously focus on incorrect anatomical regions. MedEBench sets the stage for developing more reliable and clinically effective text-guided medical image editing tools.

医学图像文本编辑可靠性评估注意力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。