arXiv:2409.01557cs.CV2024-09被引 3

TASL-Net让AI自动学习超声医生看片时的三种注意力,提升诊断精准度。

TASL-Net: Tri-Attention Selective Learning Network for Intelligent Diagnosis of Bimodal Ultrasound Video

  • 设计三重注意力机制,模拟医生看视频的时间、空间和特征关注模式。
  • 在肺、乳腺、肝脏三类数据集上准确率显著优于现有方法。
  • 适合医学影像智能诊断研究者与临床辅助决策系统开发者。

在双模态(灰度与增强型)超声视频的智能诊断中,超声医师浏览视频的方式、重点关注区域及特征对精准诊断起决定性作用。将医学知识嵌入深度学习网络不仅能提升性能,还能增强临床信心与可靠性。然而,自动捕捉视频中个体与疾病特异性的特征关注,并高效融合双模态信息仍具挑战。本文提出新型三重注意力选择学习网络(TASL-Net),在互斥变换器框架中自动嵌入超声医师的三种诊断注意力。首先,基于时间-强度曲线的视频选择器模拟医师的时间注意力,剔除冗余信息,提升计算效率;其次,提出基于结构相似性变化的早期增强位置检测器,使网络聚焦病灶内外灌注变化差异;最后,通过结合卷积与变换器的互编码策略,实现对灰度视频结构特征与增强视频灌注变化的双模态注意力。上述模块协同工作,表现优异。在肺、乳腺、肝脏三个数据集上进行了详尽实验验证。

原文摘要 · Abstract (English)

In the intelligent diagnosis of bimodal (gray-scale and contrast-enhanced) ultrasound videos, medical domain knowledge such as the way sonographers browse videos, the particular areas they emphasize, and the features they pay special attention to, plays a decisive role in facilitating precise diagnosis. Embedding medical knowledge into the deep learning network can not only enhance performance but also boost clinical confidence and reliability of the network. However, it is an intractable challenge to automatically focus on these person- and disease-specific features in videos and to enable networks to encode bimodal information comprehensively and efficiently. This paper proposes a novel Tri-Attention Selective Learning Network (TASL-Net) to tackle this challenge and automatically embed three types of diagnostic attention of sonographers into a mutual transformer framework for intelligent diagnosis of bimodal ultrasound videos. Firstly, a time-intensity-curve-based video selector is designed to mimic the temporal attention of sonographers, thus removing a large amount of redundant information while improving computational efficiency of TASL-Net. Then, to introduce the spatial attention of the sonographers for contrast-enhanced video analysis, we propose the earliest-enhanced position detector based on structural similarity variation, on which the TASL-Net is made to focus on the differences of perfusion variation inside and outside the lesion. Finally, by proposing a mutual encoding strategy that combines convolution and transformer, TASL-Net possesses bimodal attention to structure features on gray-scale videos and to perfusion variations on contrast-enhanced videos. These modules work collaboratively and contribute to superior performance. We conduct a detailed experimental validation of TASL-Net's performance on three datasets, including lung, breast, and liver.

医学影像双模态注意力机制超声诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。