arXiv:2512.17326cs.CV2025-12被引 1

开源病理图像多模态数据集与工具链,助力智能病理助手研发

Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling

  • 开发合成指令生成工具Polysome,标准化构建病理图像-文本对
  • 构建含110万条指令的公开数据集HISTAI-Instruct,覆盖2.4万张全切片图像
  • 训练出性能超越MedGemma的ANTONI-α模型,支持组织识别等精准诊断

视觉语言模型(VLMs)有望成为病理科医生的智能助手。然而,多数模型仅关注全切片图像(WSI)中的小区域,或仅提供静态的滑片级输出,且依赖非公开数据,制约了可复现性。此外,包含全切片图像与详细临床报告配对的训练数据稀缺,限制了透明、通用的VLM发展。本文提出三项主要贡献:首先,引入Polysome,一种标准化的合成指令生成工具;其次,将Polysome应用于公开的HISTAI数据集,生成包含24,259张全切片图像和超过110万条指令-响应对的HISTAI-Instruct数据集;最后,基于HISTAI-Instruct训练出ANTONI-α模型,实现全切片图像级别的视觉问答(VQA)。实验表明,ANTONI-α在组织识别、肿瘤检测和鉴别诊断任务上优于MedGemma。我们还对比了不同数据量训练的ANTONI-α变体性能。所有方法、数据与代码均公开可用。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have the potential to become co-pilots for pathologists. However, most VLMs either focus on small regions of interest within whole-slide images, provide only static slide-level outputs, or rely on data that is not publicly available, limiting reproducibility. Furthermore, training data containing WSIs paired with detailed clinical reports is scarce, restricting progress toward transparent and generalisable VLMs. We address these limitations with three main contributions. First, we introduce Polysome, a standardised tool for synthetic instruction generation. Second, we apply Polysome to the public HISTAI dataset, generating HISTAI-Instruct, a large whole-slide instruction tuning dataset spanning 24,259 slides and over 1.1 million instruction-response pairs. Finally, we use HISTAI-Instruct to train ANTONI-α, a VLM capable of visual-question answering (VQA). We show that ANTONI-α outperforms MedGemma on WSI-level VQA tasks of tissue identification, neoplasm detection, and differential diagnosis. We also compare the performance of multiple incarnations of ANTONI-α trained with different amounts of data. All methods, data, and code are publicly available.

病理分析多模态开源数据视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。