arXiv:2508.01293cs.CVcs.AI2025-08被引 1

用多智能体生成精准病理描述,提升医学图像分类效果

GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification

  • 设计多智能体系统,基于病理教材生成专业化临床描述
  • 采用描述列表代替单一提示,增强视觉与文本对齐
  • 在肾癌和肺癌数据集上达到领先水平,适合医学影像研究者

多重实例学习(MIL)是全切片图像(WSI)分类的主流方法,可高效分析千兆像素级病理切片。近期研究将视觉语言模型(VLM)引入MIL流程,通过文本描述而非简单类别名融入医学知识。然而,现有方法依赖大语言模型(LLM)生成描述或使用固定长度提示表示复杂病理概念时,受限于VLM的有限词元容量,导致编码类别信息表达力不足。此外,仅由LLM生成的描述可能缺乏领域真实性与细粒度医学特异性,影响与视觉特征的对齐效果。为此,我们提出一种视觉语言MIL框架,包含两项核心贡献:(1)基于精选病理教材与智能体专业化(如形态学、空间上下文)的领域锚定多智能体描述生成系统;(2)采用描述列表而非单个提示的文本编码策略,捕捉细粒度且互补的临床信号,实现更优视觉-文本对齐。该方法集成于VLM-MIL流程,在肾癌与肺癌数据集上表现优于单提示基线,并达到当前最优模型水平。

原文摘要 · Abstract (English)

Multiple Instance Learning (MIL) is the leading approach for whole slide image (WSI) classification, enabling efficient analysis of gigapixel pathology slides. Recent work has introduced vision-language models (VLMs) into MIL pipelines to incorporate medical knowledge through text-based class descriptions rather than simple class names. However, when these methods rely on large language models (LLMs) to generate clinical descriptions or use fixed-length prompts to represent complex pathology concepts, the limited token capacity of VLMs often constrains the expressiveness and richness of the encoded class information. Additionally, descriptions generated solely by LLMs may lack domain grounding and fine-grained medical specificity, leading to suboptimal alignment with visual features. To address these challenges, we propose a vision-language MIL framework with two key contributions: (1) A grounded multi-agent description generation system that leverages curated pathology textbooks and agent specialization (e.g., morphology, spatial context) to produce accurate and diverse clinical descriptions; (2) A text encoding strategy using a list of descriptions rather than a single prompt, capturing fine-grained and complementary clinical signals for better alignment with visual features. Integrated into a VLM-MIL pipeline, our approach shows improved performance over single-prompt class baselines and achieves results comparable to state-of-the-art models, as demonstrated on renal and lung cancer datasets.

医学图像多智能体视觉语言模型病理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。