arXiv:2505.11872cs.CV2025-05被引 6

让AI像医生一样通过解剖位置推理病灶,提升医学图像分割准确率。

PRS-Med: Position Reasoning Segmentation in Medical Imaging

  • 用视觉语言模型模拟放射科医生的定位搜索逻辑,实现结构化空间推理。
  • 在六种影像模态上,分割Dice值最高提升31.2%,显著优于现有复杂模型。
  • 专为临床设计,适合需要高可信解释性的医疗AI研发与部署场景。

基于提示的医学图像分割快速兴起,但现有方法依赖边界框等显式提示,难以有效推理对临床诊断至关重要的空间关系。通用领域模型虽尝试复杂的坐标回归,却往往缺乏医疗应用所需的结构可靠性。本文提出统一框架PRS-Med,采用临床优先的设计理念,利用集成分割解码器的医学视觉-语言模型,模仿放射科医生在特定解剖区域识别病灶的系统性“搜索模式”。为支持该推理能力,我们构建了医学位置推理分割数据集PosMed,包含跨六种成像模态的11.6万条经专家验证、具有空间锚定关系的问答对。与以往脆弱的空间推理方法不同,PosMed采用可扩展、确定性的流程,并由注册放射科医师验证,确保临床准确性。大量实验表明,基于区域的推理不仅将分割准确率(平均Dice提升达+31.2%),还提供了超越当前先进复杂推理模型的高置信度可解释性层。通过优先保障功能可靠性而非冗余技术复杂性,PRS-Med为下一代智能医疗助手提供了实用且可扩展的基准。

原文摘要 · Abstract (English)

Prompt-based medical image segmentation has rapidly emerged, yet existing methods rely on explicit prompts like bounding boxes and struggle to reason about the spatial relationships essential for clinical diagnosis. While general-domain models attempt complex coordinate regression, these approaches often lack the structured reliability required for medical applications. In this work, we introduce PRS-Med, a unified framework that adopts an elegant, clinical-first approach to position reasoning segmentation. By utilizing a medical vision-language model integrated with a segmentation decoder, PRS-Med mimics the structured "search patterns" used by radiologists to identify pathologies within specific anatomical zones. To support this robust reasoning, we present the Medical Position Reasoning Segmentation (PosMed) dataset, comprising 116,000 expert-validated, spatially grounded question-answer pairs across six imaging modalities. Unlike previous brittle attempts at spatial reasoning, PosMed leverages a scalable, deterministic pipeline validated by board-certified radiologists to ensure clinical accuracy. Extensive experiments demonstrate that our zone-based reasoning not only improves segmentation accuracy (mean Dice improvements up to +31.2\%) but also provides a high-confidence interpretability layer that outperforms state-of-the-art complex reasoning models. By prioritizing functional reliability over unnecessary technical complexity, PRS-Med offers a practical and scalable baseline for the next generation of intelligent medical assistants.

医学分割空间推理视觉语言模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。