用时空交叉注意力提升甲状腺超声动态影像的恶性判断准确率
STACT-Time: Spatio-Temporal Cross Attention for Cine Thyroid Ultrasound Time Series Classification
- 引入时空交叉注意力机制,融合超声动态序列与分割掩码特征
- 在交叉验证中达到0.91精确率和0.89 F1分数
- 适合临床辅助诊断场景,减少良性结节的过度活检
甲状腺癌是美国最常见的癌症之一。甲状腺结节常通过超声(US)成像发现,部分需进一步行细针穿刺(FNA)活检。尽管有效,FNA常导致良性结节不必要的活检,引发患者不适与焦虑。为此,美国放射学会甲状腺影像报告与数据系统(TI-RADS)被提出以减少良性活检。然而,该系统受限于观察者间差异。近期深度学习方法尝试改善风险分层,但常忽视超声动态序列提供的丰富时空上下文信息,包括多视角下的全局动态与结构变化。本文提出一种新型表示学习框架——STACT-Time,整合预训练模型生成的分割掩码与超声动态序列图像特征。通过自注意力与交叉注意力机制,模型捕捉超声动态序列中的时空上下文,并借助分割引导增强特征表示。相较于现有最优模型,本方法在交叉验证中实现0.91(±0.02)的精确率与0.89(±0.02)的F1分数,显著提升恶性肿瘤预测能力。该模型有望在保持高敏感性的同时减少良性结节的不必要活检,从而优化临床决策并改善患者预后。
原文摘要 · Abstract (English)
Thyroid cancer is among the most common cancers in the United States. Thyroid nodules are frequently detected through ultrasound (US) imaging, and some require further evaluation via fine-needle aspiration (FNA) biopsy. Despite its effectiveness, FNA often leads to unnecessary biopsies of benign nodules, causing patient discomfort and anxiety. To address this, the American College of Radiology Thyroid Imaging Reporting and Data System (TI-RADS) has been developed to reduce benign biopsies. However, such systems are limited by interobserver variability. Recent deep learning approaches have sought to improve risk stratification, but they often fail to utilize the rich temporal and spatial context provided by US cine clips, which contain dynamic global information and surrounding structural changes across various views. In this work, we propose the Spatio-Temporal Cross Attention for Cine Thyroid Ultrasound Time Series Classification (STACT-Time) model, a novel representation learning framework that integrates imaging features from US cine clips with features from segmentation masks automatically generated by a pretrained model. By leveraging self-attention and cross-attention mechanisms, our model captures the rich temporal and spatial context of US cine clips while enhancing feature representation through segmentation-guided learning. Our model improves malignancy prediction compared to state-of-the-art models, achieving a cross-validation precision of 0.91 (plus or minus 0.02) and an F1 score of 0.89 (plus or minus 0.02). By reducing unnecessary biopsies of benign nodules while maintaining high sensitivity for malignancy detection, our model has the potential to enhance clinical decision-making and improve patient outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。