用病理知识引导图像区域检索,提升癌症病理问答准确率
Path-RAG: Knowledge-Guided Key Region Retrieval for Open-ended Pathology Visual Question Answering
- 通过组织图谱技术提取病理图像中的关键区域
- 使LLaVA-Med在病理问答上准确率从38%提至47%
- 特别适合需要专家知识的长文本病理问答任务
基于病理图像的精准诊断与预后评估对癌症治疗选择至关重要。尽管深度学习在分析复杂病理图像方面取得进展,但仍因忽略组织结构和细胞组成等领域知识而表现受限。本文针对开放式的病理视觉问答(PathVQA-Open)任务,提出新框架Path-RAG,利用HistoCartography从病理图像中检索相关领域知识,显著提升性能。该方法采用以人为中心的AI策略,通过领域知识引导,精准选取病理图像中的相关补丁。实验表明,域知识引导可使LLaVA-Med在路径问答任务上的准确率从38%提升至47%,在H&E染色图像上提升达28%;对于长问答对,模型在ARCH-Open PubMed和ARCH-Open Books数据集上分别提升32.5%和30.6%。代码与数据集已公开于https://github.com/embedded-robotics/path-rag。
原文摘要 · Abstract (English)
Accurate diagnosis and prognosis assisted by pathology images are essential for cancer treatment selection and planning. Despite the recent trend of adopting deep-learning approaches for analyzing complex pathology images, they fall short as they often overlook the domain-expert understanding of tissue structure and cell composition. In this work, we focus on a challenging Open-ended Pathology VQA (PathVQA-Open) task and propose a novel framework named Path-RAG, which leverages HistoCartography to retrieve relevant domain knowledge from pathology images and significantly improves performance on PathVQA-Open. Admitting the complexity of pathology image analysis, Path-RAG adopts a human-centered AI approach by retrieving domain knowledge using HistoCartography to select the relevant patches from pathology images. Our experiments suggest that domain guidance can significantly boost the accuracy of LLaVA-Med from 38% to 47%, with a notable gain of 28% for H&E-stained pathology images in the PathVQA-Open dataset. For longer-form question and answer pairs, our model consistently achieves significant improvements of 32.5% in ARCH-Open PubMed and 30.6% in ARCH-Open Books on H\&E images. Our code and dataset is available here (https://github.com/embedded-robotics/path-rag).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。