arXiv:2606.27794cs.CV2026-06

用文本当光照调图像,让医学影像分割更准

Text as Illumination: Spatial Contrastive Retinex Learning for Language-guided Medical Image Segmentation

论文配图:Text as Illumination: Spatial Contrastive Retinex Learning for Language-guided Medical Image Segmentation
图 1 · 摘自论文原文
  • 把文本嵌入当作光照,调节图像特征增强语义一致性
  • 提出双模块与多尺度监督损失,在多个数据集上达顶尖性能
  • 适合关注跨模态对齐的医疗影像研究者

语言引导的医学图像分割(LMIS)通过融合临床文本信息,显著提升了解剖结构和病灶的分割效果。现有方法通常依赖文本与视觉特征的隐式交互或粗粒度辅助监督进行跨模态对齐,但缺乏显式的细粒度约束,导致语言与分割结果不一致。为此,本文提出文本作为光照的Retinex网络(TIRNet),将文本嵌入视为语义光照以调制特征,提升LMIS中的语义一致性。TIRNet在每个解码器阶段引入两个关键模块:(1) 基于Retinex的文本调制块(RTMB),利用正负光照图增强文本相关前景特征并抑制背景干扰;(2) 一致细节补偿块(CDCB),通过依赖光照可靠性的门控机制选择性恢复高频细节。此外,提出多尺度光照监督损失(MSIS-Loss),包含区域锚定对比损失(RGC-Loss)强制跨模态相似性集中于文本相关前景区域、抑制背景区域,以及背景抑制损失(BS-Loss)为负光照图提供像素级监督,共同确保每个解码器阶段的精确跨模态对齐。在MosMedData+和QaTa-COV19数据集上的大量实验表明,TIRNet在LMIS任务中达到当前最优性能。代码已开源:https://github.com/anaanaa/TIRNet。

原文摘要 · Abstract (English)

Language-guided Medical Image Segmentation (LMIS) has shown great potential to improve the delineation of anatomical structures and lesions by integrating clinical textual information. Existing methods generally rely on either implicit interaction between textual and visual features or auxiliary coarse-grained supervision for cross-modal alignment. However, these methods lack explicit and fine-grained constraints to ensure semantic consistency, causing a mismatch between language and the segmentation outputs. To address this issue, we propose Text-as-Illumination Retinex Network (TIRNet), a novel Retinex-inspired framework that treats text embeddings as semantic illumination for feature modulation, thereby improving semantic consistency in LMIS. TIRNet introduces two key blocks integrated at each decoder stage: (1) the Retinex-inspired Text Modulation Block (RTMB), which employs positive and negative illumination maps to enhance text-relevant foreground features and suppress background interference; and (2) the Consistent Detail Compensation Block (CDCB), which selectively recovers high-frequency details via a consistency-gated mechanism conditioned on illumination reliability. Furthermore, we propose a Multi-Scale Illumination Supervision Loss (MSIS-Loss), comprising a Region-Grounded Contrastive Loss (RGC-Loss) that enforces cross-modal similarity to be concentrated in text-relevant foreground regions and suppressed in background regions, and a Background Suppression Loss (BS-Loss) that provides pixel-level supervision for negative illumination maps, jointly ensuring a precise cross-modal alignment at each decoder stage. Extensive experiments on the MosMedData+ and QaTa-COV19 datasets demonstrate that TIRNet achieves state-of-the-art performance in LMIS. The code is available at: https://github.com/anaanaa/TIRNet.

医学图像分割跨模态对齐Retinex语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。