构建10万张高质量医疗图像编辑数据集,提升文本引导的医学影像生成能力。
MieDB-100k: A Comprehensive Dataset for Medical Image Editing
- 按感知、修改、变换三类设计编辑任务,融合理解与生成需求
- 通过专家模型与规则合成构建,经人工审核确保临床真实性
- 训练模型在多个指标上超越开源和商用模型,泛化能力强
高质量数据稀缺仍是制约多模态生成模型用于医疗图像编辑的主要瓶颈。现有医疗图像编辑数据集普遍存在多样性不足、忽视医学图像理解、难以兼顾质量与可扩展性等问题。为此,我们提出MieDB-100k,一个大规模、高质量且多样化的文本引导医疗图像编辑数据集。该数据集将编辑任务划分为感知、修改与变换三个视角,兼顾医学理解与图像生成能力。MieDB-100k通过结合模态特异性专家模型与基于规则的数据合成方法构建,并经过严格的人工审查以确保临床保真度。大量实验表明,使用MieDB-100k训练的模型在多个基准上持续优于开源与专有模型,展现出强大的泛化能力。我们期待该数据集将成为未来专业医疗图像编辑研究的基石。
原文摘要 · Abstract (English)
The scarcity of high-quality data remains a primary bottleneck in adapting multimodal generative models for medical image editing. Existing medical image editing datasets often suffer from limited diversity, neglect of medical image understanding and inability to balance quality with scalability. To address these gaps, we propose MieDB-100k, a large-scale, high-quality and diverse dataset for text-guided medical image editing. It categorizes editing tasks into perspectives of Perception, Modification and Transformation, considering both understanding and generation abilities. We construct MieDB-100k via a data curation pipeline leveraging both modality-specific expert models and rule-based data synthetic methods, followed by rigorous manual inspection to ensure clinical fidelity. Extensive experiments demonstrate that model trained with MieDB-100k consistently outperform both open-source and proprietary models while exhibiting strong generalization ability. We anticipate that this dataset will serve as a cornerstone for future advancements in specialized medical image editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。