arXiv:2511.04190cs.CV2025-11被引 1

用预训练视觉模型提取特征,构建协方差描述子提升医学图像分类性能

Covariance Descriptors Meet General Vision Encoders: Riemannian Deep Learning for Medical Image Classification

  • 从通用视觉编码器提取特征,生成协方差描述子
  • 在11个数据集上优于手工设计特征,结合DINOv2效果更优
  • 适合医学图像分析、想用几何深度学习的研究者

协方差描述子捕捉图像特征的二阶统计信息,在通用计算机视觉任务中表现优异,但在医学成像领域仍研究不足。本文探究其在传统与基于学习的医学图像分类中的有效性,重点关注专为对称正定(SPD)矩阵设计的SPDNet分类网络。我们提出从预训练通用视觉编码器(GVEs)中提取特征并构建协方差描述子,并与手工设计描述子进行对比。评估了两种GVEs——DINOv2和MedSAM——在MedMNSIT基准的11个二分类和多分类数据集上的表现。结果表明,由GVE特征生成的协方差描述子始终优于手工特征;且当SPDNet与DINOv2特征结合时,性能超越现有最优方法。研究证实,将协方差描述子与强大预训练视觉编码器结合,在医学图像分析中具有巨大潜力。

原文摘要 · Abstract (English)

Covariance descriptors capture second-order statistics of image features. They have shown strong performance in general computer vision tasks, but remain underexplored in medical imaging. We investigate their effectiveness for both conventional and learning-based medical image classification, with a particular focus on SPDNet, a classification network specifically designed for symmetric positive definite (SPD) matrices. We propose constructing covariance descriptors from features extracted by pre-trained general vision encoders (GVEs) and comparing them with handcrafted descriptors. Two GVEs - DINOv2 and MedSAM - are evaluated across eleven binary and multi-class datasets from the MedMNSIT benchmark. Our results show that covariance descriptors derived from GVE features consistently outperform those derived from handcrafted features. Moreover, SPDNet yields superior performance to state-of-the-art methods when combined with DINOv2 features. Our findings highlight the potential of combining covariance descriptors with powerful pretrained vision encoders for medical image analysis.

医学图像协方差描述子Riemannian学习视觉编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。