构建病理视觉定位新基准,实现精准区域识别
PathVG: A New Benchmark and Dataset for Pathology Visual Grounding
- 提出基于病理知识的表达-区域对齐方法
- 在33,500个标注框上达到领先性能
- 适合医学AI与多模态研究者参考
随着计算病理学快速发展,诸多AI辅助诊断任务涌现。细胞核分割虽可分类型分析,但依赖预设类别且灵活性差;病理视觉问答能实现图像级理解,却缺乏区域级检测能力。为此,我们提出新基准PathVG,旨在根据带有不同属性的语义表达检测对应区域。为评估该基准,我们构建了包含27,610张图像和33,500个语言-区域框的RefPath数据集。相较于其他领域的视觉定位,PathVG采用多尺度病理图像,并包含富含病理知识的表达。实验发现,隐含信息是主要挑战。为此,我们提出病理知识增强网络PKNet,利用大语言模型(LLMs)将含隐含信息的病理术语转化为显式视觉特征,并通过设计的知识融合模块(KFM)融合知识特征与表达特征。该方法在PathVG基准上取得最佳表现。
原文摘要 · Abstract (English)
With the rapid development of computational pathology, many AI-assisted diagnostic tasks have emerged. Cellular nuclei segmentation can segment various types of cells for downstream analysis, but it relies on predefined categories and lacks flexibility. Moreover, pathology visual question answering can perform image-level understanding but lacks region-level detection capability. To address this, we propose a new benchmark called Pathology Visual Grounding (PathVG), which aims to detect regions based on expressions with different attributes. To evaluate PathVG, we create a new dataset named RefPath which contains 27,610 images with 33,500 language-grounded boxes. Compared to visual grounding in other domains, PathVG presents pathological images at multi-scale and contains expressions with pathological knowledge. In the experimental study, we found that the biggest challenge was the implicit information underlying the pathological expressions. Based on this, we proposed Pathology Knowledge-enhanced Network (PKNet) as the baseline model for PathVG. PKNet leverages the knowledge-enhancement capabilities of Large Language Models (LLMs) to convert pathological terms with implicit information into explicit visual features, and fuses knowledge features with expression features through the designed Knowledge Fusion Module (KFM). The proposed method achieves state-of-the-art performance on the PathVG benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。