让病理AI像医生一样有条理地调用知识,提升诊断准确性。
PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs
- 构建长期记忆库,动态激活与上下文匹配的知识
- 在多个病理任务中表现最优,报告生成准确率提升超10%
- 适合需要可解释性病理诊断的临床研究与模型开发
计算病理学既需要视觉模式识别,也需动态整合结构化领域知识,如分类体系、分级标准和临床证据。实际诊断需将形态学证据与正式诊断标准关联。尽管多模态大语言模型具备强视觉-语言推理能力,却缺乏显式结构化知识整合机制与可解释的记忆控制。为此,我们受人类病理学家层次化记忆过程启发,提出PathMem——一种面向病理多模态大模型的记忆中心框架。该框架将结构化病理知识组织为长期记忆(LTM),引入记忆转换器,通过多模态记忆激活与上下文感知知识定位,实现从LTM到工作记忆(WM)的动态转化,支持下游推理中的上下文感知记忆优化。PathMem在多个基准上达到最先进性能:在WSI-Bench报告生成任务中,精确率提升12.8%,相关性提升10.1%;在开放问答诊断任务中,分别比之前基于WSI的模型提升9.7%和8.9%。
原文摘要 · Abstract (English)
Computational pathology demands both visual pattern recognition and dynamic integration of structured domain knowledge, including taxonomy, grading criteria, and clinical evidence. In practice, diagnostic reasoning requires linking morphological evidence with formal diagnostic and grading criteria. Although multimodal large language models (MLLMs) demonstrate strong vision language reasoning capabilities, they lack explicit mechanisms for structured knowledge integration and interpretable memory control. As a result, existing models struggle to consistently incorporate pathology-specific diagnostic standards during reasoning. Inspired by the hierarchical memory process of human pathologists, we propose PathMem, a memory-centric multimodal framework for pathology MLLMs. PathMem organizes structured pathology knowledge as a long-term memory (LTM) and introduces a Memory Transformer that models the dynamic transition from LTM to working memory (WM) through multimodal memory activation and context-aware knowledge grounding, enabling context-aware memory refinement for downstream reasoning. PathMem achieves SOTA performance across benchmarks, improving WSI-Bench report generation (12.8% WSI-Precision, 10.1% WSI-Relevance) and open-ended diagnosis by 9.7% and 8.9% over prior WSI-based models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。