arXiv:2506.20964cs.CVcs.AI2025-06被引 28

用多智能体系统实现病理诊断的自主推理与可解释报告生成

Evidence-based diagnostic reasoning with multi-agent copilot for human pathology

  • 基于百万级病理指令数据训练,支持文本与图像协同分析
  • 在诊断推理基准上达到高准确率,优于现有模型
  • 适合临床辅助诊断与医学AI研究者使用

病理学正经历由全切片成像和人工智能驱动的快速数字化转型。尽管基于深度学习的计算病理学已取得显著进展,但传统模型主要聚焦于图像分析,缺乏对自然语言指令和丰富文本上下文的整合。当前计算病理领域的多模态大语言模型存在训练数据不足、多图像理解支持有限及自主诊断推理能力欠缺等问题。为此,我们提出PathChat+,一种专为人类病理设计的多模态大语言模型,基于超过100万条多样化病理指令样本和近550万组问答对进行训练。在多个病理基准上的广泛评估表明,PathChat+显著优于先前的PathChat协作者,以及当前最先进的通用和病理专用模型。此外,我们提出SlideSeek,一个基于PathChat+构建的推理型多智能体系统,能够通过迭代式层级诊断推理,自主评估吉比特级全切片图像(WSIs),在DDxBench这一具有挑战性的开放式鉴别诊断基准上达到高精度,并能生成视觉相关、人类可读的摘要报告。

原文摘要 · Abstract (English)

Pathology is experiencing rapid digital transformation driven by whole-slide imaging and artificial intelligence (AI). While deep learning-based computational pathology has achieved notable success, traditional models primarily focus on image analysis without integrating natural language instruction or rich, text-based context. Current multimodal large language models (MLLMs) in computational pathology face limitations, including insufficient training data, inadequate support and evaluation for multi-image understanding, and a lack of autonomous, diagnostic reasoning capabilities. To address these limitations, we introduce PathChat+, a new MLLM specifically designed for human pathology, trained on over 1 million diverse, pathology-specific instruction samples and nearly 5.5 million question answer turns. Extensive evaluations across diverse pathology benchmarks demonstrated that PathChat+ substantially outperforms the prior PathChat copilot, as well as both state-of-the-art (SOTA) general-purpose and other pathology-specific models. Furthermore, we present SlideSeek, a reasoning-enabled multi-agent AI system leveraging PathChat+ to autonomously evaluate gigapixel whole-slide images (WSIs) through iterative, hierarchical diagnostic reasoning, reaching high accuracy on DDxBench, a challenging open-ended differential diagnosis benchmark, while also capable of generating visually grounded, humanly-interpretable summary reports.

病理诊断多智能体大语言模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。