构建细粒度眼动数据集,让肺部X光报告生成更贴近放射科医生诊断思路。
FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
- 构建首个将放射科医生眼动热图与报告逐段对齐的细粒度数据集
- 提出Gen-XAI模型,使生成报告与医生注视位置和描述高度一致
- 适合关注医学影像可解释性、临床协同智能的研究者
在计算机辅助诊断系统中,提升胸部X光(CXR)报告生成的可解释性日益重要,有助于放射科医生理解系统决策。尽管已有多种报告生成数据集与方法,但现有模型生成的报告与真实放射科医生的解读仍存在显著偏差。本研究首次提出细粒度胸部X光(FG-CXR)数据集,提供放射科医生标注的报告与对应解剖结构的眼动注意力热图的精细配对。不同于以往仅提供原始眼动序列与报告的数据集,本数据集实现了眼动位置与报告内容的精准对齐。分析表明,直接使用黑盒图像描述生成方法无法有效反映模型依赖的图像信息及关注时长。为此,我们提出一种可解释的放射科医生注意力生成网络(Gen-XAI),模拟医生诊断流程,显式约束输出与放射科医生的眼动注意力及报告内容保持一致。通过大量实验验证了方法的有效性。数据集与模型权重已公开于https://github.com/UARK-AICV/FG-CXR。
原文摘要 · Abstract (English)
Developing an interpretable system for generating reports in chest X-ray (CXR) analysis is becoming increasingly crucial in Computer-aided Diagnosis (CAD) systems, enabling radiologists to comprehend the decisions made by these systems. Despite the growth of diverse datasets and methods focusing on report generation, there remains a notable gap in how closely these models' generated reports align with the interpretations of real radiologists. In this study, we tackle this challenge by initially introducing Fine-Grained CXR (FG-CXR) dataset, which provides fine-grained paired information between the captions generated by radiologists and the corresponding gaze attention heatmaps for each anatomy. Unlike existing datasets that include a raw sequence of gaze alongside a report, with significant misalignment between gaze location and report content, our FG-CXR dataset offers a more grained alignment between gaze attention and diagnosis transcript. Furthermore, our analysis reveals that simply applying black-box image captioning methods to generate reports cannot adequately explain which information in CXR is utilized and how long needs to attend to accurately generate reports. Consequently, we propose a novel explainable radiologist's attention generator network (Gen-XAI) that mimics the diagnosis process of radiologists, explicitly constraining its output to closely align with both radiologist's gaze attention and transcript. Finally, we perform extensive experiments to illustrate the effectiveness of our method. Our datasets and checkpoint is available at https://github.com/UARK-AICV/FG-CXR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。