一个能同时处理显微图像和全切片分析的智能病理助手
A Versatile Pathology Co-pilot via Reasoning Enhanced Multimodal Large Language Model
- 用自监督方法增强推理能力,无需昂贵的思维链标注
- 支持72项任务,在230万病灶区域与18.8万全切片上表现优异
- 适合需要多任务病理分析的临床研究与AI辅助诊断系统
多模态大语言模型在计算病理学中展现出强大潜力,可整合病理图像与语言信息实现全面诊断分析。然而,现有方法受限于高昂的思维链标注成本,推理能力不足,且仅局限于病灶区域的问答任务,难以覆盖分类、检测、分割及全切片图像分类等临床需求。本文提出SmartPath-R1,一种具备强推理能力的多功能多模态模型,可同时完成病灶级与全切片级任务。通过尺度感知的监督微调与任务感知的强化学习微调,避免依赖思维链标注;采用专家混合机制实现多尺度、多任务动态分析。我们构建了包含230万病灶样本与18.8万全切片图像的大规模数据集用于训练与评估。跨72项任务的实验验证了该方法的有效性与优越性。本工作为精准病理学中多功能、强推理能力AI系统的开发迈出关键一步。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have emerged as powerful tools for computational pathology, offering unprecedented opportunities to integrate pathological images with language context for comprehensive diagnostic analysis. These models hold particular promise for automating complex tasks that traditionally require expert interpretation of pathologists. However, current MLLM approaches in pathology demonstrate significantly constrained reasoning capabilities, primarily due to their reliance on expensive chain-of-thought annotations. Additionally, existing methods remain limited to simplex application of visual question answering (VQA) at the region-of-interest (ROI) level, failing to address the full spectrum of diagnostic needs such as ROI classification, detection, segmentation, whole-slide-image (WSI) classification and VQA in clinical practice. In this study, we present SmartPath-R1, a versatile MLLM capable of simultaneously addressing both ROI-level and WSI-level tasks while demonstrating robust pathological reasoning capability. Our framework combines scale-dependent supervised fine-tuning and task-aware reinforcement fine-tuning, which circumvents the requirement for chain-of-thought supervision by leveraging the intrinsic knowledge within MLLM. Furthermore, SmartPath-R1 integrates multiscale and multitask analysis through a mixture-of-experts mechanism, enabling dynamic processing for diverse tasks. We curate a large-scale dataset comprising 2.3M ROI samples and 188K WSI samples for training and evaluation. Extensive experiments across 72 tasks validate the effectiveness and superiority of the proposed approach. This work represents a significant step toward developing versatile, reasoning-enhanced AI systems for precision pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。