arXiv:2511.00095cs.CVcs.AI2025-11被引 1

用语言交互提升脊柱CT分割准确率,支持自然语言指导修正。

SpinalSAM-R1: A Vision-Language Multimodal Interactive System for Spine CT Segmentation

  • 结合微调SAM与大模型,通过语义引导实现语言交互式分割。
  • 在脊柱图像上达到94.3%操作解析准确率,响应时间低于800毫秒。
  • 适合临床医生快速标注,降低人工标注成本,易部署使用。

从计算机断层扫描(CT)图像中分割脊柱及其邻近结构是脊柱疾病诊断与治疗的关键步骤。然而,低对比度和复杂的椎体边界限制了分割效果。尽管先进的模型如分割一切模型(SAM)在多种任务中表现良好,但在脊柱CT图像中的应用受限于高标注需求和较差的领域适应性。为此,我们提出SpinalSAM-R1,一个融合微调后的SAM与DeepSeek-R1的大模型视觉-语言交互系统,用于脊柱CT图像分割。具体地,SpinalSAM-R1引入解剖引导注意力机制以提升分割性能,并基于DeepSeek-R1设计语义驱动的交互协议,实现自然语言引导的精细化调整。系统采用低秩适配(LoRA)进行高效微调。我们在包含脊柱解剖结构的CT图像上验证了SpinalSAM-R1,实验结果表明其分割性能优越。同时,我们开发了基于PyQt5的交互软件,支持点、框和文本提示,可执行11种临床操作,解析准确率达94.3%,响应时间低于800毫秒。代码已开源于https://github.com/6jm233333/spinalsam-r1。

原文摘要 · Abstract (English)

The anatomical structure segmentation of the spine and adjacent structures from computed tomography (CT) images is a key step for spinal disease diagnosis and treatment. However, the segmentation of CT images is impeded by low contrast and complex vertebral boundaries. Although advanced models such as the Segment Anything Model (SAM) have shown promise in various segmentation tasks, their performance in spinal CT imaging is limited by high annotation requirements and poor domain adaptability. To address these limitations, we propose SpinalSAM-R1, a multimodal vision-language interactive system that integrates a fine-tuned SAM with DeepSeek-R1, for spine CT image segmentation. Specifically, our SpinalSAM-R1 introduces an anatomy-guided attention mechanism to improve spine segmentation performance, and a semantics-driven interaction protocol powered by DeepSeek-R1, enabling natural language-guided refinement. The SpinalSAM-R1 is fine-tuned using Low-Rank Adaptation (LoRA) for efficient adaptation. We validate our SpinalSAM-R1 on the spine anatomical structure with CT images. Experimental results suggest that our method achieves superior segmentation performance. Meanwhile, we develop a PyQt5-based interactive software, which supports point, box, and text-based prompts. The system supports 11 clinical operations with 94.3\% parsing accuracy and sub-800 ms response times. The software is released on https://github.com/6jm233333/spinalsam-r1.

脊柱分割多模态交互视觉语言模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。