用强化学习提升病理多模态推理能力,让模型像医生一样解释病理图像。
ReinPath: A Multimodal Reinforcement Learning Approach for Pathology
- 引入基于语义奖励的强化学习策略,提升文本描述准确性和上下文相关性。
- 在仅20%数据下训练仍优于现有方法,且在零样本分类任务中媲美CLIP。
- 构建高质量病理VQA数据集,支持复杂推理,适合医学AI研究者使用。
可解释性在计算病理学中至关重要,推动了从组织病理图像与对应文本数据中融合多模态信息的发展。然而,现有方法因缺乏支持显式推理和推断的高质量数据集,以及简单的推理过程,导致可解释性有限。为解决上述问题,我们提出一种具备强推理能力的多模态病理大语言模型。为提升生成准确且上下文相关的文本描述,设计了结合组相对策略优化的语义奖励策略。构建了一个专为复杂推理任务设计的高质量病理视觉问答(VQA)数据集。在该数据集上的全面实验表明,即使仅使用20%的数据进行训练,我们的方法仍优于当前最优模型。此外,在下游零样本图像分类任务中,性能与CLIP相当。
原文摘要 · Abstract (English)
Interpretability is significant in computational pathology, leading to the development of multimodal information integration from histopathological image and corresponding text data.However, existing multimodal methods have limited interpretability due to the lack of high-quality dataset that support explicit reasoning and inference and simple reasoning process.To address the above problems, we introduce a novel multimodal pathology large language model with strong reasoning capabilities.To improve the generation of accurate and contextually relevant textual descriptions, we design a semantic reward strategy integrated with group relative policy optimization.We construct a high-quality pathology visual question answering (VQA) dataset, specifically designed to support complex reasoning tasks.Comprehensive experiments conducted on this dataset demonstrate that our method outperforms state-of-the-art methods, even when trained with only 20% of the data.Our method also achieves comparable performance on downstream zero-shot image classification task compared with CLIP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。