arXiv:2602.19424cs.CV2026-02

用拓扑注意力分析肝癌病理切片,精准捕捉局部诊断信息。

Hepato-LLaVA: An Expert MLLM with Sparse Topo-Pack Attention for Hepatocellular Pathology Analysis on Whole Slide Images

  • 引入稀疏拓扑打包注意力,建模组织二维结构并聚合诊断证据。
  • 在3.3万条专家标注的问答数据上训练,实现肝癌诊断与描述的顶尖性能。
  • 适合医学影像分析、病理智能诊断研究者使用。

肝细胞癌诊断高度依赖于吉字节级全切片图像的解读。然而,现有计算方法受限于固定分辨率处理机制和低效特征聚合,导致严重信息丢失或高特征冗余。为此,我们提出Hepato-LLaVA,一种专用于细粒度肝细胞病理分析的多模态大语言模型。提出新颖的稀疏拓扑打包注意力机制,显式建模二维组织拓扑结构,有效将局部诊断证据聚合为语义摘要标记,同时保留全局上下文。此外,为克服多尺度数据缺乏问题,构建了基于临床场景的HepatoPathoVQA数据集,包含3.3万对经专家验证的层次化问答对。实验表明,Hepato-LLaVA在肝癌诊断与图像描述任务上达到当前最优表现。代码与实现细节见https://pris-cv.github.io/Hepto-LLaVA/。

原文摘要 · Abstract (English)

Hepatocellular Carcinoma diagnosis relies heavily on the interpretation of gigapixel Whole Slide Images. However, current computational approaches are constrained by fixed-resolution processing mechanisms and inefficient feature aggregation, which inevitably lead to either severe information loss or high feature redundancy. To address these challenges, we propose Hepato-LLaVA, a specialized Multi-modal Large Language Model designed for fine-grained hepatocellular pathology analysis. We introduce a novel Sparse Topo-Pack Attention mechanism that explicitly models 2D tissue topology. This mechanism effectively aggregates local diagnostic evidence into semantic summary tokens while preserving global context. Furthermore, to overcome the lack of multi-scale data, we present HepatoPathoVQA, a clinically grounded dataset comprising 33K hierarchically structured question-answer pairs validated by expert pathologists. Our experiments demonstrate that Hepato-LLaVA achieves state-of-the-art performance on HCC diagnosis and captioning tasks, significantly outperforming existing methods. Our code and implementation details are available at https://pris-cv.github.io/Hepto-LLaVA/.

病理分析多模态模型拓扑注意力肝癌诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。