用眼动数据对比两种模型,发现文字描述能显著提升肺部X光异常定位效果。
A Comparison of Object Detection and Phrase Grounding Models in Chest X-ray Abnormality Localization using Eye-tracking Data
- 用医生眼动数据自动标注图像区域,构建可解释性基准
- 短语定位模型性能更优(mIoU 0.36 vs. 0.20)
- 适合医疗AI可解释性研究与放射科辅助诊断系统设计
胸部疾病是全球最常见且危险的健康问题之一。目标检测与短语定位深度学习模型可解析复杂的放射科数据,辅助医生诊断。目标检测针对类别定位异常,而短语定位则针对文本描述定位异常。本文通过比较两种任务在胸部X光异常定位中的表现与可解释性,探究文本如何提升定位效果。为建立可解释性基线,我们提出一种自动管道,利用放射科医生的眼动数据生成报告句子对应的图像区域。结果表明,短语定位模型在性能(mIoU 0.36 vs. 0.20)和可解释性(包含率 0.48 vs. 0.26)上均优于目标检测模型,说明文本信息在提升胸部X光异常定位方面具有显著有效性。
原文摘要 · Abstract (English)
Chest diseases rank among the most prevalent and dangerous global health issues. Object detection and phrase grounding deep learning models interpret complex radiology data to assist healthcare professionals in diagnosis. Object detection locates abnormalities for classes, while phrase grounding locates abnormalities for textual descriptions. This paper investigates how text enhances abnormality localization in chest X-rays by comparing the performance and explainability of these two tasks. To establish an explainability baseline, we proposed an automatic pipeline to generate image regions for report sentences using radiologists' eye-tracking data. The better performance - mIoU = 0.36 vs. 0.20 - and explainability - Containment ratio 0.48 vs. 0.26 - of the phrase grounding model infers the effectiveness of text in enhancing chest X-ray abnormality localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。