医学影像模型首次实现诊断与多目标分割联合输出
RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
- 构建统一任务的放射科多模态数据集RadDiagSeg-D
- 模型可同时生成诊断文本与像素级分割掩码
- 适合临床辅助诊断系统开发人员使用
当前多数医学视觉语言模型难以在回答复杂视觉问题时同步生成诊断文本和像素级分割掩码,限制了其在临床中的应用价值。为此,我们首先提出RadDiagSeg-D数据集,将异常检测、诊断与多目标分割整合为统一的层级化任务,覆盖多种成像模态,精准支持兼具描述性文本与对应分割掩码的模型研发。基于该数据集,我们进一步提出新型视觉语言模型RadDiagSeg-M,可实现异常检测、诊断与灵活分割的联合推理。该模型输出具有高度信息量与临床实用性,有效增强了辅助诊断的上下文理解能力。最后,我们在多任务上对RadDiagSeg-M进行基准测试,验证其在多目标文本与掩码生成任务中具备强大且稳健的表现,建立了一个强有力的竞争性基线。
原文摘要 · Abstract (English)
Most current medical vision language models struggle to jointly generate diagnostic text and pixel-level segmentation masks in response to complex visual questions. This represents a major limitation towards clinical application, as assistive systems that fail to provide both modalities simultaneously offer limited value to medical practitioners. To alleviate this limitation, we first introduce RadDiagSeg-D, a dataset combining abnormality detection, diagnosis, and multi-target segmentation into a unified and hierarchical task. RadDiagSeg-D covers multiple imaging modalities and is precisely designed to support the development of models that produce descriptive text and corresponding segmentation masks in tandem. Subsequently, we leverage the dataset to propose a novel vision-language model, RadDiagSeg-M, capable of joint abnormality detection, diagnosis, and flexible segmentation. RadDiagSeg-M provides highly informative and clinically useful outputs, effectively addressing the need to enrich contextual information for assistive diagnosis. Finally, we benchmark RadDiagSeg-M and showcase its strong performance across all components involved in the task of multi-target text-and-mask generation, establishing a robust and competitive baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。