arXiv:2505.20510cs.CV2025-05NeurIPS被引 27

CPathAgent模拟病理科医生诊断流程,实现可解释的高分辨率病理图像分析。

CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic

  • 构建基于智能体的多阶段模型,按病理医生方式逐级放大观察区域。
  • 在三个图像尺度上均超越现有方法,大区域分析任务准确率显著提升。
  • 适用于需要透明诊断过程的临床场景,如医学AI辅助决策系统。

近年来计算病理学发展迅速,涌现出多种基础模型。这些模型通常依赖通用编码器结合多实例学习进行全幻灯片图像(WSI)分类,或采用多模态方法直接从图像生成报告。然而,它们无法模拟病理科医生的诊断逻辑——即先以低倍率整体扫描获得概览,再逐步聚焦可疑区域进行综合判断。现有模型直接输出最终诊断,缺乏推理过程的透明性。为此,我们提出CPathAgent,一种创新的基于智能体的方法,通过根据视觉特征自主导航于全幻灯片图像之间,生成更透明、可解释的诊断摘要。我们设计了多阶段训练策略,统一了局部块、区域和全图级别的理解能力,以复现病理医生跨尺度认知与推理过程。此外,我们构建了首个专家验证的大区域分析基准数据集PathMMU-HR2,填补了局部块与全幻灯片之间的关键临床尺度空白——病理医生通常会重点检查数个关键大区域而非整张切片。大量实验证明,CPathAgent在三个不同图像尺度的基准测试中持续优于现有方法,验证了其智能体式诊断方法的有效性,为计算病理学指明了可解释性新方向。

原文摘要 · Abstract (English)

Recent advances in computational pathology have led to the emergence of numerous foundation models. These models typically rely on general-purpose encoders with multi-instance learning for whole slide image (WSI) classification or apply multimodal approaches to generate reports directly from images. However, these models cannot emulate the diagnostic approach of pathologists, who systematically examine slides at low magnification to obtain an overview before progressively zooming in on suspicious regions to formulate comprehensive diagnoses. Instead, existing models directly output final diagnoses without revealing the underlying reasoning process. To address this gap, we introduce CPathAgent, an innovative agent-based approach that mimics pathologists' diagnostic workflow by autonomously navigating across WSI based on observed visual features, thereby generating substantially more transparent and interpretable diagnostic summaries. To achieve this, we develop a multi-stage training strategy that unifies patch-level, region-level, and WSI-level capabilities within a single model, which is essential for replicating how pathologists understand and reason across diverse image scales. Additionally, we construct PathMMU-HR2, the first expert-validated benchmark for large region analysis. This represents a critical intermediate scale between patches and whole slides, reflecting a key clinical reality where pathologists typically examine several key large regions rather than entire slides at once. Extensive experiments demonstrate that CPathAgent consistently outperforms existing approaches across benchmarks at three different image scales, validating the effectiveness of our agent-based diagnostic approach and highlighting a promising direction for computational pathology.

病理图像可解释性智能体多尺度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。