arXiv:2509.06105cs.CV2025-09EMNLP被引 2

提升病理图像与报告的跨模态理解能力,解决细微差异识别难题。

PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology

  • 构建分层语义理解评测基准,聚焦病理报告的复合推理能力。
  • 在7个病理数据集上实现领先性能,显著优于现有模型。
  • 专为病理场景设计训练方案,增强模型对细微结构差异的感知。

准确分析病理图像对自动化肿瘤诊断至关重要,但因组织图像结构高度相似且形态变化细微,仍具挑战性。现有视觉-语言(VL)模型难以捕捉复杂推理需求以解读结构化病理报告。为此,我们提出PathoHR-Bench,一个用于评估VL模型在病理领域中分层语义理解与组合推理能力的新基准。评测结果表明,现有模型无法有效建模复杂的跨模态关系,限制了其在临床中的应用。为此,我们进一步引入一种针对病理领域的VL训练方案,通过生成增强与扰动样本,支持多模态对比学习。实验验证显示,该方法在PathoHR-Bench及六个额外病理数据集上均达到当前最优性能,证明其在细粒度病理表征上的有效性。

原文摘要 · Abstract (English)

Accurate analysis of pathological images is essential for automated tumor diagnosis but remains challenging due to high structural similarity and subtle morphological variations in tissue images. Current vision-language (VL) models often struggle to capture the complex reasoning required for interpreting structured pathological reports. To address these limitations, we propose PathoHR-Bench, a novel benchmark designed to evaluate VL models' abilities in hierarchical semantic understanding and compositional reasoning within the pathology domain. Results of this benchmark reveal that existing VL models fail to effectively model intricate cross-modal relationships, hence limiting their applicability in clinical setting. To overcome this, we further introduce a pathology-specific VL training scheme that generates enhanced and perturbed samples for multimodal contrastive learning. Experimental evaluations demonstrate that our approach achieves state-of-the-art performance on PathoHR-Bench and six additional pathology datasets, highlighting its effectiveness in fine-grained pathology representation.

病理分析视觉语言模型细粒度识别跨模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。