arXiv:2507.01029cs.LGcs.AI2025-07被引 4

让AI像病理专家一样一步步推理,减少错误判断。

PathCoT: Chain-of-Thought Prompting for Zero-shot Pathology Visual Reasoning

  • 用病理知识引导大模型逐步推理,提升专业性。
  • 在PathMMU数据集上准确率显著优于现有方法。
  • 适合医疗AI、医学影像分析领域的研究者使用。

随着生成式人工智能和指令微调技术的发展,多模态大语言模型(MLLM)在通用推理任务中取得了显著进展。得益于思维链(CoT)方法,MLLM能够分步解决视觉推理问题。然而,在病理视觉推理任务中,现有MLLM仍面临两大挑战:(1)因缺乏领域专业知识导致模型幻觉,表现不佳;(2)CoT中的额外推理步骤可能引入误差,造成答案发散。为此,我们提出PathCoT,一种新颖的零样本思维链提示方法,将病理专家知识融入MLLM的推理过程,并引入自评估机制以缓解答案发散。具体而言,PathCoT通过先验知识引导MLLM以病理专家视角进行图像分析,结合领域专长完成完整推理。同时,该方法包含自评估环节,对直接输出与CoT生成的结果进行评估,最终确定可靠答案。在PathMMU数据集上的实验表明,该方法在病理视觉理解与推理任务中具有显著有效性。

原文摘要 · Abstract (English)

With the development of generative artificial intelligence and instruction tuning techniques, multimodal large language models (MLLMs) have made impressive progress on general reasoning tasks. Benefiting from the chain-of-thought (CoT) methodology, MLLMs can solve the visual reasoning problem step-by-step. However, existing MLLMs still face significant challenges when applied to pathology visual reasoning tasks: (1) LLMs often underperforms because they lack domain-specific information, which can lead to model hallucinations. (2) The additional reasoning steps in CoT may introduce errors, leading to the divergence of answers. To address these limitations, we propose PathCoT, a novel zero-shot CoT prompting method which integrates the pathology expert-knowledge into the reasoning process of MLLMs and incorporates self-evaluation to mitigate divergence of answers. Specifically, PathCoT guides the MLLM with prior knowledge to perform as pathology experts, and provides comprehensive analysis of the image with their domain-specific knowledge. By incorporating the experts' knowledge, PathCoT can obtain the answers with CoT reasoning. Furthermore, PathCoT incorporates a self-evaluation step that assesses both the results generated directly by MLLMs and those derived through CoT, finally determining the reliable answer. The experimental results on the PathMMU dataset demonstrate the effectiveness of our method on pathology visual understanding and reasoning.

病理推理思维链多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。