arXiv:2603.19957cs.CVcs.AI2026-03被引 1

让病理报告生成更结构化,准确率超74%且安全率近97%

HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction

  • 分层视觉语言对齐框架,用三模块处理多图编码与结构化输出
  • 在74.9万例真实数据上实现74.7%临床可接受准确率,安全率97.3%
  • 适合医学AI研发者,支持跨医院部署且性能稳定

病理报告是包含诊断结论、组织学分级及附加检测结果的结构化多粒度文档,覆盖一个或多个解剖部位;现有病理视觉-语言模型(VLM)将其输出简化为扁平标签或自由文本。本文提出轻量级VLM框架HiPath,基于冻结的UNI2和Qwen3主干网络,以结构化报告生成为主要训练目标。三个可训练模块共1500万参数,分别解决:多图像视觉编码的分层补丁聚合(HiPA)、基于最优传输的跨模态对齐(HiCL),以及基于槽位的掩码诊断生成(Slot-MDP)。在来自三家医院的74.9万例真实中国病理病例上训练,HiPath达到68.9%严格准确率和74.7%临床可接受准确率,安全率达97.3%,优于所有基线方法。跨医院评估显示仅下降3.4个百分点严格准确率,安全率维持97.1%。

原文摘要 · Abstract (English)

Pathology reports are structured, multi-granular documents encoding diagnostic conclusions, histological grades, and ancillary test results across one or more anatomical sites; yet existing pathology vision-language models (VLMs) reduce this output to a flat label or free-form text. We present HiPath, a lightweight VLM framework built on frozen UNI2 and Qwen3 backbones that treats structured report prediction as its primary training objective. Three trainable modules totalling 15M parameters address complementary aspects of the problem: a Hierarchical Patch Aggregator (HiPA) for multi-image visual encoding, Hierarchical Contrastive Learning (HiCL) for cross-modal alignment via optimal transport, and Slot-based Masked Diagnosis Prediction (Slot-MDP) for structured diagnosis generation. Trained on 749K real-world Chinese pathology cases from three hospitals, HiPath achieves 68.9% strict and 74.7% clinically acceptable accuracy with a 97.3% safety rate, outperforming all baselines under the same frozen backbone. Cross-hospital evaluation confirms generalisation with only a 3.4pp drop in strict accuracy while maintaining 97.1% safety.

病理报告视觉语言模型结构化生成医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。