arXiv:2511.06665cs.CVcs.AI2025-11AAAI被引 1

提出医学诊断分割新任务,结合视觉语言模型实现精准病灶定位与可解释诊断。

Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks

  • 设计区域感知的视觉-语言相似性掩码模块,提升跨模态对齐精度
  • 在多疾病数据集上实现分割与诊断双项性能超越基线
  • 适合需要可解释性医疗辅助决策的研究者与临床应用

尽管像素级医学图像分析取得进展,现有分割模型很少联合考虑医学分割与诊断任务。然而,为患者提供可解释的诊断结果至关重要。本文提出医学诊断分割(MDS)新任务,旨在理解临床查询并生成对应分割掩码与诊断结论。为此,我们构建了包含多样化多模态多疾病医学图像及其分割掩码和诊断思维链的M3DS数据集,通过自动化诊断思维链生成流程创建。同时,提出Sim4Seg框架,利用区域感知的视觉-语言相似性掩码(RVLS2M)模块提升诊断分割性能,并探索测试时扩展策略以增强表现。实验表明,该方法在分割与诊断两项任务上均优于基线。

原文摘要 · Abstract (English)

Despite significant progress in pixel-level medical image analysis, existing medical image segmentation models rarely explore medical segmentation and diagnosis tasks jointly. However, it is crucial for patients that models can provide explainable diagnoses along with medical segmentation results. In this paper, we introduce a medical vision-language task named Medical Diagnosis Segmentation (MDS), which aims to understand clinical queries for medical images and generate the corresponding segmentation masks as well as diagnostic results. To facilitate this task, we first present the Multimodal Multi-disease Medical Diagnosis Segmentation (M3DS) dataset, containing diverse multimodal multi-disease medical images paired with their corresponding segmentation masks and diagnosis chain-of-thought, created via an automated diagnosis chain-of-thought generation pipeline. Moreover, we propose Sim4Seg, a novel framework that improves the performance of diagnosis segmentation by taking advantage of the Region-Aware Vision-Language Similarity to Mask (RVLS2M) module. To improve overall performance, we investigate a test-time scaling strategy for MDS tasks. Experimental results demonstrate that our method outperforms the baselines in both segmentation and diagnosis.

医学图像视觉语言可解释诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。