arXiv:2508.15960cs.CV2025-08被引 3

用视觉语言模型精准识别肾小球疾病亚型,少样本下仍达高准确率

Glo-VLMs: Leveraging Vision-Language Models for Fine-Grained Diseased Glomerulus Classification

  • 结合病理图像与临床文本提示,实现图文联合表征学习
  • 8样本/类时准确率达0.7416,宏AUC达0.9045
  • 适用于数据稀缺的精细医学图像分类任务

视觉语言模型(VLMs)在数字病理学中展现出巨大潜力,但在区分肾小球亚型等细粒度、疾病特异性分类任务中效果有限。由于亚型间形态差异细微,且视觉模式与精确临床术语对齐困难,肾脏病理自动诊断极具挑战。本文提出Glo-VLMs,一个系统性框架,用于在数据受限条件下适配预训练VLMs进行细粒度肾小球分类。方法结合精选病理图像与临床文本提示,促进对细微肾病亚型的联合表征学习。通过在少样本学习范式下评估多种VLM架构与适配策略,研究方法选择与标注数据量对临床相关场景性能的影响。所有模型均使用标准化多分类指标进行公平比较,旨在明确大模型在专业临床研究中的实际需求与潜力。结果表明,微调VLMs在每类仅8个样本的情况下,达到0.7416的准确率、0.9045的宏AUC和0.5277的F1分数,证明即使在高度受限的监督条件下,基础模型仍可有效适配于细粒度医学图像分类。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have shown considerable potential in digital pathology, yet their effectiveness remains limited for fine-grained, disease-specific classification tasks such as distinguishing between glomerular subtypes. The subtle morphological variations among these subtypes, combined with the difficulty of aligning visual patterns with precise clinical terminology, make automated diagnosis in renal pathology particularly challenging. In this work, we explore how large pretrained VLMs can be effectively adapted to perform fine-grained glomerular classification, even in scenarios where only a small number of labeled examples are available. In this work, we introduce Glo-VLMs, a systematic framework designed to explore the adaptation of VLMs to fine-grained glomerular classification in data-constrained settings. Our approach leverages curated pathology images alongside clinical text prompts to facilitate joint image-text representation learning for nuanced renal pathology subtypes. By assessing various VLMs architectures and adaptation strategies under a few-shot learning paradigm, we explore how both the choice of method and the amount of labeled data impact model performance in clinically relevant scenarios. To ensure a fair comparison, we evaluate all models using standardized multi-class metrics, aiming to clarify the practical requirements and potential of large pretrained models for specialized clinical research applications. As a result, fine-tuning the VLMs achieved 0.7416 accuracy, 0.9045 macro-AUC, and 0.5277 F1-score with only 8 shots per class, demonstrating that even with highly limited supervision, foundation models can be effectively adapted for fine-grained medical image classification.

医学图像视觉语言模型少样本学习肾病分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。