用文字提示实现医学影像序列的精准目标分割与跟踪
Text-Promptable Propagation for Referring Medical Image Sequence Segmentation
- 通过跨模态交互识别文本描述的目标,用Transformer实现连续跟踪
- 在涵盖4种模态20类器官的基准上,分割精度显著优于现有方法
- 适合医疗影像智能分析、临床辅助诊断等场景使用
指代式医学影像序列分割(Ref-MISS)是一项新兴且具有挑战性的任务,旨在根据自然语言描述对医学影像序列(如内窥镜、超声、CT和MRI)中的解剖结构进行分割。该任务具有重要的临床应用价值,能提升医学影像解读的用户友好性。现有2D和3D分割模型难以显式追踪影像序列中的目标,且缺乏交互式文本引导能力。为此,我们提出文本提示传播模型(TPP),用于指代式医学影像序列分割。TPP捕捉序列图像与其对应文本描述之间的内在关联,通过跨模态指代交互识别目标,并利用基于Transformer的三重传播机制,以文本嵌入为查询实现序列间的连续跟踪。为支持该任务,我们构建了一个大规模基准数据集Ref-MISS-Bench,覆盖4种成像模态和20种不同器官及病灶。在该基准上的实验结果表明,TPP在医学分割和指代视频目标分割任务中均持续优于现有最先进方法。
原文摘要 · Abstract (English)
Referring Medical Image Sequence Segmentation (Ref-MISS) is a novel and challenging task that aims to segment anatomical structures in medical image sequences (\emph{e.g.} endoscopy, ultrasound, CT, and MRI) based on natural language descriptions. This task holds significant clinical potential and offers a user-friendly advancement in medical imaging interpretation. Existing 2D and 3D segmentation models struggle to explicitly track objects of interest across medical image sequences, and lack support for nteractive, text-driven guidance. To address these limitations, we propose Text-Promptable Propagation (TPP), a model designed for referring medical image sequence segmentation. TPP captures the intrinsic relationships among sequential images along with their associated textual descriptions. Specifically, it enables the recognition of referred objects through cross-modal referring interaction, and maintains continuous tracking across the sequence via Transformer-based triple propagation, using text embeddings as queries. To support this task, we curate a large-scale benchmark, Ref-MISS-Bench, which covers 4 imaging modalities and 20 different organs and lesions. Experimental results on this benchmark demonstrate that TPP consistently outperforms state-of-the-art methods in both medical segmentation and referring video object segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。