模仿医生阅片习惯,用新模型提升肺部CT多病灶识别准确率
Imitating Radiological Scrolling: A Global-Local Attention Model for 3D Chest CT Volumes Multi-Label Anomaly Classification
- 设计全局-局部注意力机制,模拟医生逐层翻看CT的阅读方式
- 在两个公开数据集上实现优于传统CNN和Transformer的分类性能
- 适合需要精准定位与上下文理解的医学影像分析场景
CT扫描数量激增带来巨大工作压力,亟需自动化工具辅助放射科医生完成器官分割、异常分类和报告生成。三维CT多标签分类因数据体量大、病灶类型多样而极具挑战。现有基于卷积神经网络的方法难以捕捉长程依赖,视觉变换器则需大量预训练,实用性受限。此外,这些方法未显式建模医生在浏览切片时的导航行为,该过程需兼顾全局上下文与局部细节。本文提出CT-Scroll,一种专为模拟医生阅片流程设计的全局-局部注意力模型。在两个公开数据集上的实验表明,该模型有效提升了分类性能,并通过消融实验证明了各组件的贡献。
原文摘要 · Abstract (English)
The rapid increase in the number of Computed Tomography (CT) scan examinations has created an urgent need for automated tools, such as organ segmentation, anomaly classification, and report generation, to assist radiologists with their growing workload. Multi-label classification of Three-Dimensional (3D) CT scans is a challenging task due to the volumetric nature of the data and the variety of anomalies to be detected. Existing deep learning methods based on Convolutional Neural Networks (CNNs) struggle to capture long-range dependencies effectively, while Vision Transformers require extensive pre-training, posing challenges for practical use. Additionally, these existing methods do not explicitly model the radiologist's navigational behavior while scrolling through CT scan slices, which requires both global context understanding and local detail awareness. In this study, we present CT-Scroll, a novel global-local attention model specifically designed to emulate the scrolling behavior of radiologists during the analysis of 3D CT scans. Our approach is evaluated on two public datasets, demonstrating its efficacy through comprehensive experiments and an ablation study that highlights the contribution of each model component.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。