用语义引导定位病灶,提升医学影像分割与诊断的准确性。
Rad-VLSM: A Cross-Modal Framework with Semantics-Assisted Prompting for Medical Segmentation and Diagnosis

- 通过视觉-语言对齐定位病灶候选区域,生成提示框
- 多区域聚合策略提升分割稳定性,准确率超基线3.2%
- 融合影像与影像组学特征,实现病灶级诊断依据
医学图像分割若能支持诊断,则更具临床价值。然而,诊断相关病灶线索往往细微且局部,现有模型常被背景组织、声学伪影及无关视觉关联干扰。为此,我们提出Rad-VLSM,一种两阶段跨模态框架,实现语义引导的病灶聚焦、鲁棒分割与视觉化诊断。第一阶段基于BLIP-2的视觉-语言对齐模块,在语义引导下识别病灶相关候选区域,并转换为框提示;第二阶段将提示输入基于SAM的多任务网络,采用多候选区域聚合策略提升提示稳定性并指导分割。预测掩码作为空间先验用于诊断,视觉-影像组学融合头整合病灶感知视觉特征与选定影像组学描述符。通过语义信息进行定位而非直接预测,Rad-VLSM降低文本到诊断的依赖,使诊断建立在病灶级证据之上。在私有临床乳腺超声数据集及公开基准上实验表明,Rad-VLSM在分割与诊断性能上表现优异,具有良好泛化能力。
原文摘要 · Abstract (English)
Medical image segmentation is more clinically valuable when it supports diagnosis rather than merely producing lesion masks. However, diagnostically relevant lesion cues are often subtle and localized, while existing models may be distracted by background tissues, acoustic artifacts, and irrelevant visual correlations. To address this problem, we propose Rad-VLSM, a two-stage cross-modal framework for semantics-assisted lesion focusing, robust segmentation, and visually grounded diagnosis. In the first stage, a BLIP-2-based vision-language alignment module identifies lesion-related candidate regions under semantic guidance and converts them into box prompts. In the second stage, these prompts are fed into a SAM-based multitask network, where a multi-candidate region aggregation strategy improves prompt stability and guides lesion segmentation. The predicted masks are then used as spatial priors for diagnosis, and a visual-radiomics fusion head integrates lesion-aware visual features with selected radiomics descriptors. By using semantic information for localization rather than direct prediction, Rad-VLSM reduces text-to-diagnosis dependence and grounds diagnosis in lesion-level evidence. Experiments on a private clinical breast ultrasound dataset and public benchmarks show that Rad-VLSM achieves strong segmentation and diagnostic performance with favorable generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。