arXiv:2509.25594cs.CVcs.AI2025-09被引 1

K-Prism统一医学图像分割,融合三类知识实现灵活高效分割。

K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model

  • 用双提示结构编码语义先验、上下文示例和用户交互信息。
  • 在18个数据集上达到当前最优,跨模态、跨任务表现稳定。
  • 适合需要多源知识融合的医疗影像分析场景。

医学图像分割是临床决策的基础,但现有模型仍呈碎片化:通常仅基于单一知识来源,且针对特定任务、模态或器官。这与临床实践不符——医生会自然融合解剖先验、参考病例推理和实时交互反馈。我们提出K-Prism,一种统一分割框架,系统集成三种知识范式:(i) 从标注数据中学到的语义先验,(ii) 从少量参考样本中获取的上下文知识,(iii) 用户输入(如点击或涂画)带来的交互反馈。核心思想是将异构知识编码为双提示表示:1维稀疏提示定义待分割内容,2维密集提示指示关注位置,并通过混合专家(MoE)解码器动态路由。该设计支持范式间灵活切换及跨任务联合训练,无需修改架构。在涵盖CT、MRI、X-ray、病理、超声等模态的18个公开数据集上的实验证明,K-Prism在语义、上下文和交互分割设置下均达到领先性能。

原文摘要 · Abstract (English)

Medical image segmentation is fundamental to clinical decision-making, yet existing models remain fragmented. They are usually trained on single knowledge sources and specific to individual tasks, modalities, or organs. This fragmentation contrasts sharply with clinical practice, where experts seamlessly integrate diverse knowledge: anatomical priors from training, exemplar-based reasoning from reference cases, and iterative refinement through real-time interaction. We present $\textbf{K-Prism}$, a unified segmentation framework that mirrors this clinical flexibility by systematically integrating three knowledge paradigms: (i) $\textit{semantic priors}$ learned from annotated datasets, (ii) $\textit{in-context knowledge}$ from few-shot reference examples, and (iii) $\textit{interactive feedback}$ from user inputs like clicks or scribbles. Our key insight is that these heterogeneous knowledge sources can be encoded into a dual-prompt representation: 1-D sparse prompts defining $\textit{what}$ to segment and 2-D dense prompts indicating $\textit{where}$ to attend, which are then dynamically routed through a Mixture-of-Experts (MoE) decoder. This design enables flexible switching between paradigms and joint training across diverse tasks without architectural modifications. Comprehensive experiments on 18 public datasets spanning diverse modalities (CT, MRI, X-ray, pathology, ultrasound, etc.) demonstrate that K-Prism achieves state-of-the-art performance across semantic, in-context, and interactive segmentation settings.

医学图像分割知识融合多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。