用专家看片行为训练可解释的病理诊断智能体
Pathology-CoT: Learning Visual Chain-of-Thought Agent from Expert Whole Slide Image Diagnosis Behavior
- 从医生看片日志中自动提取行为指令与理由,构建可扩展的标注数据集
- 在胃肠道淋巴结转移检测上实现100%召回率,超越现有最强模型
- 适合需要可解释性与临床落地的病理AI研发团队使用
全切片图像诊断是一个交互式、多阶段的过程,涉及放大倍数切换和视野移动。尽管近期病理基础模型表现优异,但能自主决定观察区域、调节放大倍数并生成可解释诊断的实用智能体系统仍匮乏。这一瓶颈主要源于数据:缺乏大规模、贴近临床的专家看片行为标注,这些行为依赖经验且未被记录在教材或网络中,因此未被大语言模型训练覆盖。本文提出一种框架,通过三大突破解决此问题:首先,AI会话记录器无缝集成于标准全切片图像查看器,无感记录常规操作,并将查看器日志转化为标准化的行为命令与边界框;其次,轻量级人机协同审核将AI生成的决策理由转化为Pathology-CoT数据集,包含“看哪里”与“为何重要”的配对信息,使标注速度提升六倍;基于该数据,构建两阶段智能体Pathology-o3,先定位关键区域,再进行行为引导推理。在斯坦福医学内部验证集上达到100%召回率,瑞典独立外部验证集上达97.6%召回率,优于当前最先进的OpenAI o3模型,并具备跨骨干网络泛化能力。据我们所知,Pathology-CoT是病理学领域首批基于行为的智能体系统。将日常查看日志转化为可扩展、专家验证的监督信号,使病理智能体成为现实,并为可对齐人类、可升级的临床人工智能开辟路径。
原文摘要 · Abstract (English)
Diagnosing a whole-slide image is an interactive, multi-stage process of changing magnification and moving between fields. Although recent pathology foundation models demonstrated superior performances, practical agentic systems that decide what field to examine next, adjust magnification, and deliver explainable diagnoses are still lacking. Such limitation is largely bottlenecked by data: scalable, clinically aligned supervision of expert viewing behavior that is tacit and experience-based, not documented in textbooks or internet, and therefore absent from LLM training. Here we introduce a framework designed to address this challenge through three key breakthroughs. First, the AI Session Recorder seamlessly integrates with standard whole-slide image viewers to unobtrusively record routine navigation and convert the viewer logs into standardized behavioral commands and bounding boxes. Second, a lightweight human-in-the-loop review turns AI-drafted rationales for behavioral commands into the Pathology-CoT dataset, a form of paired "where to look" and "why it matters", enabling six-fold faster labeling compared to manual constructing such Chain-of-Thought dataset. Using this behavioral data, we build Pathology-o3, a two-stage agent that first proposes important ROIs and then performs behavior-guided reasoning. On the gastrointestinal lymph-node metastasis detection task, our method achieved 100 recall on the internal validation from Stanford Medicine and 97.6 recall on an independent external validation from Sweden, exceeding the state-of-the-art OpenAI o3 model and generalizing across backbones. To our knowledge, Pathology-CoT constitutes one of the first behavior-grounded agentic systems in pathology. Turning everyday viewer logs into scalable, expert-validated supervision, our framework makes agentic pathology practical and establishes a path to human-aligned, upgradeable clinical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。