用大模型模拟病理医生看片过程,让诊断结果可解释。
PathAgent: Toward Interpretable Analysis of Whole-slide Pathology Images via Large Language Model-based Agentic Reasoning
- 基于大模型构建自主分析代理,分步定位病灶区域
- 零样本跨数据集表现优于专用模型,支持开放问答
- 生成自然语言推理链,适合临床辅助与医学教育
全切片图像(WSI)分析需要迭代、基于证据的推理过程,类似于病理医生动态放大、重新聚焦并自我修正以收集证据。然而现有计算流程通常缺乏明确的推理轨迹,导致预测结果不透明且难以解释。为此,我们提出PathAgent,一种无需训练的大语言模型(LLM)代理框架,模仿人类专家的反思式、分步分析方法。PathAgent可自主探索WSI,通过导航模块精确识别显著微区域,利用感知模块提取形态学视觉线索,并将这些发现整合到执行器中持续演化的自然语言推理链中。整个观察与决策序列形成明确的思维链,实现完全可解释的预测。在五个具有挑战性的数据集上评估,PathAgent展现出强大的零样本泛化能力,在开放式与约束型视觉问答任务中均超越任务特定基线。此外,与人类病理医生的协作评估确认了PathAgent作为透明且临床可信的诊断助手的潜力。
原文摘要 · Abstract (English)
Analyzing whole-slide images (WSIs) requires an iterative, evidence-driven reasoning process that parallels how pathologists dynamically zoom, refocus, and self-correct while collecting the evidence. However, existing computational pipelines often lack this explicit reasoning trajectory, resulting in inherently opaque and unjustifiable predictions. To bridge this gap, we present PathAgent, a training-free, large language model (LLM)-based agent framework that emulates the reflective, stepwise analytical approach of human experts. PathAgent can autonomously explore WSI, iteratively and precisely locating significant micro-regions using the Navigator module, extracting morphology visual cues using the Perceptor, and integrating these findings into the continuously evolving natural language trajectories in the Executor. The entire sequence of observations and decisions forms an explicit chain-of-thought, yielding fully interpretable predictions. Evaluated across five challenging datasets, PathAgent exhibits strong zero-shot generalization, surpassing task-specific baselines in both open-ended and constrained visual question-answering tasks. Moreover, a collaborative evaluation with human pathologists confirms PathAgent's promise as a transparent and clinically grounded diagnostic assistant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。