用语言重对齐视觉特征,提升跨物种病理图像识别能力
Lost in Translation: How Language Re-Aligns Vision for Cross-Species Pathology
- 引入语义锚定机制,通过语言稳定视觉特征空间
- 跨物种识别AUC达66.31%,比基准提升5.67%
- 揭示了物种主导对齐导致的语义坍缩新问题
基础模型在计算病理学中的应用日益广泛,但其在跨癌症与跨物种迁移下的行为尚不明确。本研究考察了微调CPath-CLIP在人与犬类组织切片上的相同癌种、跨癌种及跨物种条件下的癌症检测表现,使用受试者工作特征曲线下面积(AUC)评估性能。少样本微调使同癌种性能从64.9%提升至72.6% AUC,跨癌种从56.84%提升至66.31% AUC。跨物种评估显示,尽管组织匹配可实现有意义迁移,但性能仍低于当前最优基准(H-optimus-0: 84.97% AUC),表明标准视觉-语言对齐对跨物种泛化效果不佳。嵌入空间分析显示肿瘤与正常原型间余弦相似度超0.99。Grad-CAM显示原型模型仍受限于领域,而语言引导模型关注保守肿瘤形态。为此提出语义锚定,利用语言为视觉特征提供稳定坐标系。消融实验表明收益源于文本对齐机制本身,而非文本编码器复杂性。对比H-optimus-0发现,CPath-CLIP失败源于内在嵌入坍缩,而文本对齐有效规避该问题。同癌种与跨癌种分类分别获得8.52%和5.67%提升。识别出一种此前未被描述的失效模式:由物种主导对齐引发的语义坍缩,而非视觉信息缺失。结果表明,语言可作为控制机制,实现语义重解释而无需重新训练。
原文摘要 · Abstract (English)
Foundation models are increasingly applied to computational pathology, yet their behavior under cross-cancer and cross-species transfer remains unspecified. This study investigated how fine-tuning CPath-CLIP affects cancer detection under same-cancer, cross-cancer, and cross-species conditions using whole-slide image patches from canine and human histopathology. Performance was measured using area under the receiver operating characteristic curve (AUC). Few-shot fine-tuning improved same-cancer (64.9% to 72.6% AUC) and cross-cancer performance (56.84% to 66.31% AUC). Cross-species evaluation revealed that while tissue matching enables meaningful transfer, performance remains below state-of-the-art benchmarks (H-optimus-0: 84.97% AUC), indicating that standard vision-language alignment is suboptimal for cross-species generalization. Embedding space analysis revealed extremely high cosine similarity (greater than 0.99) between tumor and normal prototypes. Grad-CAM shows prototype-based models remain domain-locked, while language-guided models attend to conserved tumor morphology. To address this, we introduce Semantic Anchoring, which uses language to provide a stable coordinate system for visual features. Ablation studies reveal that benefits stem from the text-alignment mechanism itself, regardless of text encoder complexity. Benchmarking against H-optimus-0 shows that CPath-CLIP's failure stems from intrinsic embedding collapse, which text alignment effectively circumvents. Additional gains were observed in same-cancer (8.52%) and cross-cancer classification (5.67%). We identified a previously uncharacterized failure mode: semantic collapse driven by species-dominated alignment rather than missing visual information. These results demonstrate that language acts as a control mechanism, enabling semantic re-interpretation without retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。