arXiv:2507.11055cs.CV2025-07ICCV被引 6

用原型替代文本,让医学图像分割无需配对报告也能精准分割。

Alleviating Textual Reliance in Medical Language-guided Segmentation via Prototype-driven Semantic Approximation

  • 构建原型空间,从文本中提取关键语义用于图像引导。
  • 在无文本情况下仍能实现接近有文本的分割精度。
  • 适合缺乏报告配对数据或需先分割后报告的临床场景。

医学语言引导分割通过结合临床文本报告提升图像分割效果,但其依赖成对图文数据存在两大问题:一是大量仅有图像的数据无法用于训练;二是推理受限于已有报告,难以应用于分割先行的临床场景。为此,本文提出首个无文本依赖的原型驱动学习框架ProLearn。核心是原型驱动语义近似(PSA)模块,通过提炼文本中的分割相关语义建立离散紧凑的原型空间,支持无需文本输入时的查询响应机制,从而近似提供语义引导。在QaTa-COV19、MosMedData+和Kvasir-SEG数据集上的实验表明,当文本资源有限时,ProLearn优于现有先进方法。

原文摘要 · Abstract (English)

Medical language-guided segmentation, integrating textual clinical reports as auxiliary guidance to enhance image segmentation, has demonstrated significant improvements over unimodal approaches. However, its inherent reliance on paired image-text input, which we refer to as ``textual reliance", presents two fundamental limitations: 1) many medical segmentation datasets lack paired reports, leaving a substantial portion of image-only data underutilized for training; and 2) inference is limited to retrospective analysis of cases with paired reports, limiting its applicability in most clinical scenarios where segmentation typically precedes reporting. To address these limitations, we propose ProLearn, the first Prototype-driven Learning framework for language-guided segmentation that fundamentally alleviates textual reliance. At its core, we introduce a novel Prototype-driven Semantic Approximation (PSA) module to enable approximation of semantic guidance from textual input. PSA initializes a discrete and compact prototype space by distilling segmentation-relevant semantics from textual reports. Once initialized, it supports a query-and-respond mechanism which approximates semantic guidance for images without textual input, thereby alleviating textual reliance. Extensive experiments on QaTa-COV19, MosMedData+ and Kvasir-SEG demonstrate that ProLearn outperforms state-of-the-art language-guided methods when limited text is available.

医学分割少文本原型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。