用视觉语言模型实现少量标注下的肾小球精细分型,提升病理诊断精度。
From Classification to Cross-Modal Understanding: Leveraging Vision-Language Models for Fine-Grained Renal Pathology
- 将肾小球亚型分类设为少样本任务,评估专用与通用视觉语言模型表现。
- 仅4-8个样本/类型时,专用模型已能有效区分亚型,性能显著优于基础模型。
- 模型选择与微调策略共同影响诊断效果,适合临床数据稀缺场景使用。
细粒度肾小球亚型分型是肾活检解读的核心,但临床上有价值标签稀少且难获取。现有计算病理方法多在全监督下进行粗略疾病分类,依赖图像模型,尚不清楚视觉语言模型(VLMs)如何在数据受限条件下适配有意义的亚型分型。本文将细粒度肾小球亚型分类建模为临床真实的少样本问题,系统评估病理专用与通用视觉语言模型在此设定下的表现。不仅考察分类性能(准确率、AUC、F1),还分析学习表征的几何特性,包括图像与文本嵌入的对齐程度及亚型可分性。通过联合分析样本数、模型架构、领域知识与适配策略,为真实临床数据约束下的模型选择与训练提供指导。结果表明,病理专用视觉语言主干网络搭配基础微调,是最有效的起点。即使每亚型仅4-8个标注样本,这些模型已能捕捉差异,显著提升判别力与校准性,额外监督仍带来增量改善。同时发现,正负样本判别能力与图文对齐同等重要。总体而言,监督水平与适配策略共同塑造诊断性能与多模态结构,为模型选择、适配策略与标注投入提供依据。
原文摘要 · Abstract (English)
Fine-grained glomerular subtyping is central to kidney biopsy interpretation, but clinically valuable labels are scarce and difficult to obtain. Existing computational pathology approaches instead tend to evaluate coarse diseased classification under full supervision with image-only models, so it remains unclear how vision-language models (VLMs) should be adapted for clinically meaningful subtyping under data constraints. In this work, we model fine-grained glomerular subtyping as a clinically realistic few-shot problem and systematically evaluate both pathology-specialized and general-purpose vision-language models under this setting. We assess not only classification performance (accuracy, AUC, F1) but also the geometry of the learned representations, examining feature alignment between image and text embeddings and the separability of glomerular subtypes. By jointly analyzing shot count, model architecture and domain knowledge, and adaptation strategy, this study provides guidance for future model selection and training under real clinical data constraints. Our results indicate that pathology-specialized vision-language backbones, when paired with the vanilla fine-tuning, are the most effective starting point. Even with only 4-8 labeled examples per glomeruli subtype, these models begin to capture distinctions and show substantial gains in discrimination and calibration, though additional supervision continues to yield incremental improvements. We also find that the discrimination between positive and negative examples is as important as image-text alignment. Overall, our results show that supervision level and adaptation strategy jointly shape both diagnostic performance and multimodal structure, providing guidance for model selection, adaptation strategies, and annotation investment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。