构建真实病理图像退化数据集,提升视觉语言模型鲁棒性。
Histopath-C: Towards Realistic Domain Shifts for Histopathology Vision-Language Adaptation
- 动态合成病理图像退化效果,模拟真实世界分布偏移。
- 提出LATTE方法,在多个病理数据集上超越现有测试时自适应技术。
- 适合关注医学图像模型泛化能力的研究者与临床应用开发者。
医学视觉语言模型(VLM)在组织病理学等医学影像领域表现优异,依赖预训练对比模型融合视觉与文本信息。然而,病理图像常面临染色差异、污染、模糊和噪声等严重域偏移,显著降低模型下游性能。本文提出Histopath-C基准,通过真实合成的退化扰动模拟数字病理学中的分布偏移。该框架可动态对任意数据集施加退化,并实时评估测试时自适应(TTA)机制。我们进一步提出LATTE——一种基于多文本模板的归纳式低秩适配策略,缓解病理学VLM对不同文本输入的敏感性。实验表明,该方法在多种病理数据集上优于为自然图像设计的前沿TTA方法,验证了其在病理图像中实现稳健适应的有效性。代码与数据见https://github.com/Mehrdad-Noori/Histopath-C。
原文摘要 · Abstract (English)
Medical Vision-language models (VLMs) have shown remarkable performances in various medical imaging domains such as histo\-pathology by leveraging pre-trained, contrastive models that exploit visual and textual information. However, histopathology images may exhibit severe domain shifts, such as staining, contamination, blurring, and noise, which may severely degrade the VLM's downstream performance. In this work, we introduce Histopath-C, a new benchmark with realistic synthetic corruptions designed to mimic real-world distribution shifts observed in digital histopathology. Our framework dynamically applies corruptions to any available dataset and evaluates Test-Time Adaptation (TTA) mechanisms on the fly. We then propose LATTE, a transductive, low-rank adaptation strategy that exploits multiple text templates, mitigating the sensitivity of histopathology VLMs to diverse text inputs. Our approach outperforms state-of-the-art TTA methods originally designed for natural images across a breadth of histopathology datasets, demonstrating the effectiveness of our proposed design for robust adaptation in histopathology images. Code and data are available at https://github.com/Mehrdad-Noori/Histopath-C.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。