arXiv:2505.04192cs.CVcs.AI2025-05被引 1

首个整合病理视频与诊断推理的多模态模型,助力智能病理分析。

ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos

  • 融合单张切片、自动分割视频与人工标注视频,模拟真实诊断流程。
  • 基于4278对视频-诊断链式思考数据训练,实现详细描述与确诊输出。
  • 适合医学AI研究者和临床辅助系统开发者使用。

我们提出ViDRiP-LLaVA,首个在计算病理学领域集成三类图像场景的大型多模态模型(LMM),包括单张切片图像、自动分割的病理视频片段及人工标注的病理视频。该设计紧密贴合病理学家的实际诊断过程。通过生成详细的组织学描述并最终给出明确的诊断报告,ViDRiP-LLaVA实现了视觉叙事与诊断推理的深度融合。核心是构建了包含4278个视频与诊断相关链式思考指令对的ViDRiP-Instruct数据集,数据来源于YouTube上的教学病理视频。尽管高质量数据对提升诊断推理至关重要,但其制作耗时且数量有限。为此,我们采用从已有单图指令数据集迁移知识的方法,在弱标注的关键帧片段上进行训练,再在人工标注视频上微调。ViDRiP-LLaVA为病理视频分析建立了新基准,并为未来支持临床决策的AI系统提供了坚实基础。代码、数据与模型已公开:https://github.com/QuIIL/ViDRiP-LLaVA。

原文摘要 · Abstract (English)

We present ViDRiP-LLaVA, the first large multimodal model (LMM) in computational pathology that integrates three distinct image scenarios, including single patch images, automatically segmented pathology video clips, and manually segmented pathology videos. This integration closely mirrors the natural diagnostic process of pathologists. By generating detailed histological descriptions and culminating in a definitive sign-out diagnosis, ViDRiP-LLaVA bridges visual narratives with diagnostic reasoning. Central to our approach is the ViDRiP-Instruct dataset, comprising 4278 video and diagnosis-specific chain-of-thought instructional pairs sourced from educational histopathology videos on YouTube. Although high-quality data is critical for enhancing diagnostic reasoning, its creation is time-intensive and limited in volume. To overcome this challenge, we transfer knowledge from existing single-image instruction datasets to train on weakly annotated, keyframe-extracted clips, followed by fine-tuning on manually segmented videos. ViDRiP-LLaVA establishes a new benchmark in pathology video analysis and offers a promising foundation for future AI systems that support clinical decision-making through integrated visual and diagnostic reasoning. Our code, data, and model are publicly available at: https://github.com/QuIIL/ViDRiP-LLaVA.

病理分析多模态诊断推理视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。