arXiv:2512.23545cs.CVcs.AI2025-12被引 1

让病理诊断模型像医生一样主动找证据,提升诊断准确性。

PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis

  • 构建智能代理框架,分三阶段推进诊断:初判、找证据、定结论。
  • 在多场景下诊断准确率超越现有模型,尤其擅长发现微小病灶特征。
  • 适合临床辅助诊断系统研发者及计算病理研究者参考。

近期病理基础模型显著提升了视觉表征学习与多模态交互能力。然而,多数模型仍采用静态推理范式,仅对全切片图像进行一次处理即生成预测,缺乏对模糊诊断的重新评估或针对性证据获取。这与临床诊断中通过反复观察切片并请求进一步检查来修正假设的流程相悖。我们提出 PathFound,一种支持证据寻求推理的智能体式多模态模型。PathFound 融合病理视觉基础模型、视觉语言模型与强化学习训练的推理模型,通过初始诊断、证据获取、最终决策三个阶段,实现主动信息采集与诊断优化。在多个大型多模态模型中,采用该策略均一致提升诊断准确率,证明了证据寻求工作流在计算病理中的有效性。其中,PathFound 在多样临床场景中达到当前最优表现,展现出发现细微特征(如核异型性、局部浸润)的强大潜力。

原文摘要 · Abstract (English)

Recent pathological foundation models have substantially advanced visual representation learning and multimodal interaction. However, most models still rely on a static inference paradigm in which whole-slide images are processed once to produce predictions, without reassessment or targeted evidence acquisition under ambiguous diagnoses. This contrasts with clinical diagnostic workflows that refine hypotheses through repeated slide observations and further examination requests. We propose PathFound, an agentic multimodal model designed to support evidence-seeking inference in pathological diagnosis. PathFound integrates the power of pathological visual foundation models, vision-language models, and reasoning models trained with reinforcement learning to perform proactive information acquisition and diagnosis refinement by progressing through the initial diagnosis, evidence-seeking, and final decision stages. Across several large multimodal models, adopting this strategy consistently improves diagnostic accuracy, indicating the effectiveness of evidence-seeking workflows in computational pathology. Among these models, PathFound achieves state-of-the-art diagnostic performance across diverse clinical scenarios and demonstrates strong potential to discover subtle details, such as nuclear features and local invasions.

病理诊断智能代理多模态证据寻求

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。