arXiv:2608.13939cs.CVcs.AI2026-08

用文本描述对齐超声图像,实现甲状腺结节细粒度分类

CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification

  • 将TI-RADS特征文本嵌入与超声图像嵌入对齐,通过对比损失提升表征能力
  • 在数据有限且不平衡情况下,分类准确率超越主流方法5.2个百分点
  • 适合医学影像多模态学习、甲状腺结节智能诊断的研究者参考

超声是评估甲状腺结节的主要影像手段,ACR TI-RADS框架通过五类超声特征归纳为五个风险等级(TR1-TR5)。尽管临床广泛应用,多数深度学习方法仍集中于二分类恶性与否,而多类别预测与特征级监督的利用仍不充分,主要受限于标注数据稀少。本研究构建了包含600个甲状腺结节的STN数据集,涵盖横切面与纵切面超声图像、边界框标注及全部五类TI-RADS特征的完整标签。遵循临床决策流程,探究结构化特征信息如何指导训练阶段的表征学习,同时推理时仅需图像输入。实验表明,基于标准化特征描述生成的文本嵌入可作为TI-RADS风险等级的稳定代理表示。据此提出CMCNet,通过中心-间隔对比损失将图像嵌入对齐至固定文本嵌入,同步增强类内紧凑性与类间分离性。结果表明,该对齐策略比直接多任务学习更具数据效率和鲁棒性,在不平衡场景下显著优于InfoNCE、中心损失、强基线多任务模型及VQA型多模态模型。数据集公开于doi: 10.5281/zenodo.19125693,源码见:https://www.healthinformaticslab.org/supp/

原文摘要 · Abstract (English)

Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5). Although widely adopted in clinical practice, most deep learning approaches focus on binary malignancy classification, while multi-class prediction and explicit utilization of feature-level supervision remain underexplored, largely due to limited annotated data. In this study, we introduce the STN dataset of 600 thyroid nodules with paired transverse and longitudinal ultrasound images, bounding box annotations, and complete labels for all five TI-RADS feature categories. Following the clinical decision process, we investigate how structured feature information can guide representation learning during training while requiring only images at inference. We demonstrate that text embeddings derived from standardized feature descriptions form a stable surrogate representation for TI-RADS risk levels. Based on this observation, we propose CMCNet, which aligns image embeddings to fixed textual embeddings via a Center-Margin Contrastive Loss that simultaneously promotes intra-class compactness and inter-class separation. Experimental results show that this embedding alignment strategy is more data-efficient and robust than direct multitask learning, and consistently outperforms InfoNCE, center loss, a strong multitask baseline, and a VQA-style multimodal model, particularly in imbalanced settings. The dataset is freely available at doi: 10.5281/zenodo.19125693 and the source code is available at: https://www.healthinformaticslab.org/supp/.

医学影像多模态学习甲状腺结节对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。