arXiv:2501.08042cs.CVcs.AI2025-01被引 1

用视觉语言模型提升尤文肉瘤诊断准确率,效果优于传统方法。

Exploring visual language models as a powerful tool in the diagnosis of Ewing Sarcoma

  • 采用视觉语言预训练策略,结合多实例学习框架
  • 诊断准确率显著提升,参数量与计算成本大幅降低
  • 适合病理医生辅助诊断,尤其对青少年患者有参考价值

尤文肉瘤(ES)以高密度小圆蓝细胞、无结构组织为特征,主要影响10至19岁青少年,是重大健康问题。基于人工智能的组织病理图像自动分析系统有望提升ES诊断准确性。本研究首次探索不同预训练策略在区分ES与其他形态相似的软组织或骨肉瘤方面的特征提取能力,采用数字化组织微阵列数据集,在多实例学习范式下对比了视觉-语言监督(VLS)与全监督ImageNet预训练。结果表明,使用领域内数据进行VLS适配后,诊断准确率显著提升;且模型不仅提高了分类预测精度,还大幅减少可训练参数量与计算开销。

原文摘要 · Abstract (English)

Ewing's sarcoma (ES), characterized by a high density of small round blue cells without structural organization, presents a significant health concern, particularly among adolescents aged 10 to 19. Artificial intelligence-based systems for automated analysis of histopathological images are promising to contribute to an accurate diagnosis of ES. In this context, this study explores the feature extraction ability of different pre-training strategies for distinguishing ES from other soft tissue or bone sarcomas with similar morphology in digitized tissue microarrays for the first time, as far as we know. Vision-language supervision (VLS) is compared to fully-supervised ImageNet pre-training within a multiple instance learning paradigm. Our findings indicate a substantial improvement in diagnostic accuracy with the adaption of VLS using an in-domain dataset. Notably, these models not only enhance the accuracy of predicted classes but also drastically reduce the number of trainable parameters and computational costs.

病理诊断视觉语言模型尤文肉瘤多实例学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。