用简单指令实现胸部X光片病灶分割,数据集超百万对。
Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset
- 基于用户指令进行病灶分割,无需专业术语输入。
- 构建110万对指令-答案数据集,覆盖7类病灶。
- 模型可生成分割结果与文本解释,适合临床辅助使用。
当前胸部X光片(CXRs)病灶分割模型受限于目标标签数量少且依赖复杂专家级文本输入,难以实用。为此,本文提出指令引导病灶分割(ILS),作为引用图像分割(RIS)在医学领域的适配,可根据简单用户指令分割多种病灶类型。我们构建了MIMIC-ILS——首个大规模的胸部X光病灶分割指令-答案数据集,通过全自动多模态流水线从原始影像和报告中生成标注。该数据集包含192,000张图像、91,000个唯一分割掩码及110万条指令-答案对,覆盖7类主要病灶类型。为验证其有效性,我们训练了ROSALIA模型,该模型在MIMIC-ILS上微调后,能精准分割多种病灶并提供文本解释。实验表明,该模型在新提出的任务中表现优异,验证了本方法与数据集的有效性。数据集与模型已开源。
原文摘要 · Abstract (English)
The applicability of current lesion segmentation models for chest X-rays (CXRs) has been limited both by a small number of target labels and the reliance on complex, expert-level text inputs, creating a barrier to practical use. To address these limitations, we introduce instruction-guided lesion segmentation (ILS), a medical-domain adaptation of referring image segmentation (RIS) designed to segment diverse lesion types based on simple, user-friendly instructions. Under this task, we construct MIMIC-ILS, the first large-scale instruction-answer dataset for CXR lesion segmentation, using our fully automated multimodal pipeline that generates annotations from CXR images and their corresponding reports. MIMIC-ILS contains 1.1M instruction-answer pairs derived from 192K images and 91K unique segmentation masks, covering seven major lesion types. To empirically demonstrate its utility, we present ROSALIA, a LISA model fine-tuned on the MIMIC-ILS dataset. ROSALIA can segment diverse lesions and provide textual explanations in response to user instructions. The model achieves high accuracy in our newly proposed task, highlighting the effectiveness of our pipeline and the value of MIMIC-ILS as a foundational resource for pixel-level CXR lesion grounding. The dataset and model are available at https://github.com/checkoneee/ROSALIA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。