用图像标签+文本提示,让AI自动分割眼底OCT病变,省去人工标注。
Text-Driven Weakly Supervised OCT Lesion Segmentation with Structural Guidance
- 结合结构信息与文本描述生成像素级伪标签。
- 在三个OCT数据集上达到顶尖分割效果。
- 适合医疗影像标注成本高的场景。
精准分割光学相干断层扫描(OCT)图像对诊断和监测视网膜疾病至关重要。然而,像素级标注的高劳动成本限制了监督学习在大规模数据集上的应用。弱监督语义分割(WSSS)通过使用图像级标签等较弱监督形式,有望降低标注负担。尽管如此,弱监督本身信息有限。本文提出一种仅需图像级标签的新型WSSS框架,融合结构引导与文本驱动的双重指导,生成高质量像素级伪标签。该框架包含两个视觉处理模块:一个处理原始OCT图像,另一个处理加入异常信号的分层分割图,使模型能将病变关联到对应解剖层。同时,借助大规模预训练模型提供两类文本引导:基于标签的局部语义描述,以及不依赖领域的合成描述,虽以自然图像术语表达,却捕捉空间与关系语义,有助于生成全局一致表征。通过多模态融合视觉与文本特征,方法实现语义与结构的相关性对齐,显著提升病变定位与分割性能。在三个OCT数据集上的实验表明,本方法达到当前最优水平,展现出推动医学影像诊断准确率与效率的潜力。
原文摘要 · Abstract (English)
Accurate segmentation of Optical Coherence Tomography (OCT) images is crucial for diagnosing and monitoring retinal diseases. However, the labor-intensive nature of pixel-level annotation limits the scalability of supervised learning for large datasets. Weakly Supervised Semantic Segmentation (WSSS) offers a promising alternative by using weaker forms of supervision, such as image-level labels, to reduce the annotation burden. Despite its advantages, weak supervision inherently carries limited information. We propose a novel WSSS framework with only image-level labels for OCT lesion segmentation that integrates structural and text-driven guidance to produce high-quality, pixel-level pseudo labels. The framework employs two visual processing modules: one that processes the original OCT images and another that operates on layer segmentations augmented with anomalous signals, enabling the model to associate lesions with their corresponding anatomical layers. Complementing these visual cues, we leverage large-scale pretrained models to provide two forms of textual guidance: label-derived descriptions that encode local semantics, and domain-agnostic synthetic descriptions that, although expressed in natural image terms, capture spatial and relational semantics useful for generating globally consistent representations. By fusing these visual and textual features in a multi-modal framework, our method aligns semantic meaning with structural relevance, thereby improving lesion localization and segmentation performance. Experiments on three OCT datasets demonstrate state-of-the-art results, highlighting its potential to advance diagnostic accuracy and efficiency in medical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。