arXiv:2604.17570cs.CVcs.AI2026-04中稿 · CVPR被引 1

首个专为血涂片设计的多层级视觉语言框架与评测基准

PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation

论文配图:PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation
图 1 · 摘自论文原文
  • 构建首个血涂片图文数据集PBSInstr,含353幅全切片图像和29000个细胞级标注
  • 提出专用模型PBS-VL,在六项任务中超越通用病理模型性能
  • 适用于血液病理智能辅助诊断,助力临床决策支持系统研发

外周血涂片(PBS)是血液病理学中关键的显微检查手段,生成全切片图像(WSI)。与实体组织病理不同,其诊断重点在于单个细胞形态而非组织结构,因此在视觉特征和诊断逻辑上具有独特性。然而,现有用于病理学的多模态大语言模型主要基于实体组织的全切片图像开发,难以泛化至血涂片场景。为此,我们构建了首个针对血涂片解释的视觉语言数据集PBSInstr,包含353幅血涂片全切片图像及对应的显微印象段落,以及29,000个细胞级图像裁剪,均带有细胞类型标签与形态描述。同时,该数据集还包含27,000个细胞图像问答对和1,286个全片问答对,以支持指令微调。在此基础上,我们开发了面向血液病理的视觉语言模型PBS-VL,可实现细胞与全片级别的多层级解释。为进一步全面评估血涂片理解能力,我们构建了包含四类问题与六项任务的PBSBench视觉问答基准。实验表明,PBS-VL显著优于现有通用及病理领域多模态模型,验证了专用数据的价值。我们已开源代码、数据集与模型权重,推动后续研究。本框架为开发实际可用的血液病理智能辅助系统奠定了基础。

原文摘要 · Abstract (English)

Peripheral Blood Smear (PBS) is a critical microscopic examination in hematopathology that yields whole-slide imaging (WSI). Unlike solid tissue pathology, PBS interpretation focuses on individual cell morphologies rather than tissue architecture, making it distinct in both visual characteristics and diagnostic reasoning. However, current multimodal large language models (MLLMs) for pathology are primarily developed on solid-tissue WSIs and struggle to generalize to PBS. To bridge this gap, we construct PBSInstr, the first vision-language dataset for PBS interpretation, comprising 353 PBS WSIs paired with microscopic impression paragraphs and 29k cell-level image crops annotated with cell type labels and morphological descriptions. To facilitate instruction tuning, PBSInstr further includes 27k question-answer (QA) pairs for cell crops and 1,286 QA pairs for PBS slides. Building upon PBSInstr, we develop PBS-VL, a hematopathology-tailored vision-language model for multi-level PBS interpretation at both cell and slide levels. To comprehensively evaluate PBS understanding, we construct PBSBench, a visual question answering (VQA) benchmark featuring four question categories and six PBS interpretation tasks. Experiments show that PBS-VL outperforms existing general-purpose and pathology MLLMs, underscoring the value of PBS-specific data. We release our code, datasets, and model weights to facilitate future research. Our proposed framework lays the foundation for developing practical AI assistants supporting decision-making in hematopathology.

医学影像视觉语言血涂片多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。