用强化学习训练视觉语言模型,提升胃肠道病理诊断的准确性与可解释性。
DiagR1: A Vision-Language Model Trained via Reinforcement Learning for Digestive Pathology Diagnosis
- 通过提示论证策略融合病变分类与解剖位置信息,增强图像特征捕捉
- 在真实病理报告生成任务中,临床相关性提升18.7%,结构完整性提高32.4%
- 适用于需要高可信度诊断辅助的临床场景,尤其适合病理医生协作
多模态大模型在自动化病理图像分析中展现出巨大潜力。然而,当前用于胃肠道病理的多模态模型受限于数据质量与推理透明度:公开数据集普遍存在噪声和标注不全,导致视觉语言模型在生成诊断文本时易出现事实性幻觉;同时缺乏明确的中间推理链,使输出难以审计,临床信任度低。为此,我们构建了一个大规模胃肠道病理数据集,包含显微描述与诊断结论,并提出一种融合病变分类与解剖部位信息的提示论证策略,引导模型更好捕捉图像特异性特征并保持生成语义一致性。此外,采用结合监督微调与组相对策略优化(GRPO)的后训练流程,提升推理质量与输出结构。在真实世界病理报告生成任务上的实验表明,该方法在生成质量、结构完整性和临床相关性上显著优于开源及商用基线模型。相比现有方案,本方法临床相关性提升18.7%,结构完整性提高32.4%,诊断错误减少41.2%,展现出更高的准确性和临床实用性。
原文摘要 · Abstract (English)
Multimodal large models have shown great potential in automating pathology image analysis. However, current multimodal models for gastrointestinal pathology are constrained by both data quality and reasoning transparency: pervasive noise and incomplete annotations in public datasets predispose vision language models to factual hallucinations when generating diagnostic text, while the absence of explicit intermediate reasoning chains renders the outputs difficult to audit and thus less trustworthy in clinical practice. To address these issues, we construct a large scale gastrointestinal pathology dataset containing both microscopic descriptions and diagnostic conclusions, and propose a prompt argumentation strategy that incorporates lesion classification and anatomical site information. This design guides the model to better capture image specific features and maintain semantic consistency in generation. Furthermore, we employ a post training pipeline that combines supervised fine tuning with Group Relative Policy Optimization (GRPO) to improve reasoning quality and output structure. Experimental results on real world pathology report generation tasks demonstrate that our approach significantly outperforms state of the art open source and proprietary baselines in terms of generation quality, structural completeness, and clinical relevance. Our solution outperforms state of the art models with 18.7% higher clinical relevance, 32.4% improved structural completeness, and 41.2% fewer diagnostic errors, demonstrating superior accuracy and clinical utility compared to existing solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。