arXiv:2501.12524eess.IVcs.AI2025-01中稿 · IEEE ISBI 2025被引 3

用自蒸馏视觉变压器提升小样本肺超声严重程度评分

Efficient Lung Ultrasound Severity Scoring Using Dedicated Feature Extractor

  • 自蒸馏预训练ViT+双层VLAD特征聚合,减少对标注数据依赖
  • 微调后在帧级和视频级评分上超越全监督方法
  • 适合医疗影像少样本场景,可解释性强

新冠疫情推动超声成像在疾病检测中的应用,因其无创、低成本和便携性。然而,公开的超声数据集规模小且标注不充分,制约了AI模型的训练。本文提出MeDiVLAD,一种针对多层级肺超声(LUS)严重程度评分的新方法。通过自知识蒸馏预训练视觉变换器(ViT)并无需标签,再采用双层VLAD聚合帧级特征。实验表明,经少量微调后,MeDiVLAD在帧级和视频级评分任务中均优于传统全监督方法,并具备高质量分类推理能力。该性能支持关键病理区域自动识别,为更广泛的医学视频分类任务提供稳健解决方案。

原文摘要 · Abstract (English)

With the advent of the COVID-19 pandemic, ultrasound imaging has emerged as a promising technique for COVID-19 detection, due to its non-invasive nature, affordability, and portability. In response, researchers have focused on developing AI-based scoring systems to provide real-time diagnostic support. However, the limited size and lack of proper annotation in publicly available ultrasound datasets pose significant challenges for training a robust AI model. This paper proposes MeDiVLAD, a novel pipeline to address the above issue for multi-level lung-ultrasound (LUS) severity scoring. In particular, we leverage self-knowledge distillation to pretrain a vision transformer (ViT) without label and aggregate frame-level features via dual-level VLAD aggregation. We show that with minimal finetuning, MeDiVLAD outperforms conventional fully-supervised methods in both frame- and video-level scoring, while offering classification reasoning with exceptional quality. This superior performance enables key applications such as the automatic identification of critical lung pathology areas and provides a robust solution for broader medical video classification tasks.

肺超声视觉变压器少样本学习医疗影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。