arXiv:2508.15256cs.CV2025-08ICCV被引 2

用病理知识增强视觉语言模型,提升罕见病变检测准确率。

Normal and Abnormal Pathology Knowledge-Augmented Vision-Language Model for Anomaly Detection in Pathology Images

  • 基于预训练视觉语言模型,加入正常与异常病理知识
  • 在两个淋巴结数据集上达到当前最优检测与定位性能
  • 通过图文关联实现可解释性,适合医学影像分析场景

计算病理学中的异常检测旨在识别罕见且数据稀缺的病变,但现有方法多为工业场景设计,受限于计算资源、组织结构多样性和可解释性不足。为此,我们提出 Ano-NAViLa:一种融合正常与异常病理知识的视觉语言模型,用于病理图像异常检测。该模型基于预训练视觉语言模型,仅引入轻量级可训练MLP,通过整合病理知识显著提升对病理图像变异性的鲁棒性与检测精度,并借助图像-文本关联实现可解释性。在来自不同器官的两个淋巴结数据集上评估,Ano-NAViLa 在异常检测与定位任务中均优于现有模型,达到当前最优水平。

原文摘要 · Abstract (English)

Anomaly detection in computational pathology aims to identify rare and scarce anomalies where disease-related data are often limited or missing. Existing anomaly detection methods, primarily designed for industrial settings, face limitations in pathology due to computational constraints, diverse tissue structures, and lack of interpretability. To address these challenges, we propose Ano-NAViLa, a Normal and Abnormal pathology knowledge-augmented Vision-Language model for Anomaly detection in pathology images. Ano-NAViLa is built on a pre-trained vision-language model with a lightweight trainable MLP. By incorporating both normal and abnormal pathology knowledge, Ano-NAViLa enhances accuracy and robustness to variability in pathology images and provides interpretability through image-text associations. Evaluated on two lymph node datasets from different organs, Ano-NAViLa achieves the state-of-the-art performance in anomaly detection and localization, outperforming competing models.

异常检测病理图像视觉语言模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。