用报告自动生成医学图像分割,支持新器官无需重训。
ReportMedSAM: Guiding Segmentation Through Radiology Reports

- 用可学习的概念库替代固定关键词提取,动态匹配报告与分割模块。
- 在AbdomenAtlas 3.0上达到可比分割精度,对同义词变化鲁棒。
- 新器官加入无需重训,适合医疗多任务扩展场景。
自由格式的放射科报告包含丰富的临床描述,但将其用于可靠分割仍具挑战,源于自然语言固有的变异性。现有方法常依赖预定义器官短语或脆弱的规则式推理时提取,限制其对新解剖结构的扩展性,并易受语言变化影响。为此,我们提出ReportMedSAM,一种以报告驱动的框架,将离散提取替换为可学习的概念库。通过冻结的医学视觉-语言编码器(BiomedCLIP),我们利用对比学习将器官级概念嵌入与大规模临床语料对齐,建立互正交的语义锚点。该方法显式缓解器官级语义坍缩,确保对多样化临床同义词(如“renal”与“kidney”)的高度鲁棒性。推理时,临床报告被嵌入并匹配至概念库,动态激活特定任务的Mixture-of-Experts(MoE)模块。这种解耦设计允许新增概念与专家,而无需重训已有组件,实现参数隔离扩展,同时保持已学专家不变。在AbdomenAtlas 3.0数据集上的评估表明,ReportMedSAM能有效解析自由格式报告,实现竞争性分割精度,并无缝、无干扰地扩展至新临床任务。
原文摘要 · Abstract (English)
Free-form radiology reports contain rich clinical descriptions, yet converting them for reliable segmentation remains challenging due to the inherent variability of natural language. Existing pipelines often rely on predefined organ phrases or brittle rule-based inference-time extraction, which limits their scalability to novel anatomical structures and makes them sensitive to linguistic variations. To address this, we propose ReportMedSAM, a report-driven framework that replaces discrete extraction with a learnable concept bank. By leveraging a frozen medical vision-language encoder (BiomedCLIP), we align organ-level concept embeddings with large-scale clinical corpora through contrastive learning, establishing mutually orthogonal semantic anchors. Our approach explicitly mitigates organ-level semantic collapse and ensures high robustness against diverse clinical synonyms (e.g., "renal" vs. "kidney" ). During inference, a clinical report is embedded and matched against this concept bank to dynamically activate task-specific Mixture-of-Experts (MoE) modules. This decoupled design allows new concepts and experts to be added without retraining existing components, providing a parameter-isolated extension mechanism while keeping previously learned experts unchanged. Evaluated on the AbdomenAtlas 3.0 dataset, ReportMedSAM effectively interprets free-form reports, achieves competitive segmentation accuracy, and demonstrates seamless, non-interfering extension to novel clinical tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。