arXiv:2512.03445cs.CVcs.AI2025-12被引 2

用多智能体生成高质量医学图文数据,提升皮肤病识别效果

Multi-Aspect Knowledge-Enhanced Medical Vision-Language Pretraining with Multi-Agent Data Generation

  • 用智能体系统自动合成带知识的图文描述,提升数据质量
  • 将长文本拆解为多个知识维度,实现图文细粒度对齐
  • 在8个数据集上达成顶尖零样本性能,适合医学多模态研究

视觉-语言预训练(VLP)在医学图像分析中展现出强大潜力,能利用大规模无标注图文对进行表征学习。然而现有方法面临网络收集数据噪声多、长篇非结构化医学文本难处理的问题。为此,本文提出集成多智能体数据生成(MAGEN)与基于本体的多方面知识增强(O-MAKE)的VLP框架。MAGEN通过基础模型辅助的描述生成与检索验证流程,提升数据质量;O-MAKE将复杂长文本分解为不同知识维度,实现全局与局部像素级的细粒度对齐,并借助本体引导机制显式建模医学概念关系。我们在皮肤科领域验证了该框架的有效性,实验表明其在疾病分类和跨模态检索任务上于8个数据集上均达到当前最优零样本表现。代码与扩充数据集Derm1M-AgentAug(含超40万张皮肤图像-文本对)将开源于https://github.com/SiyuanYan1/Derm1M。

原文摘要 · Abstract (English)

Vision-language pretraining (VLP) has emerged as a powerful paradigm in medical image analysis, enabling representation learning from large-scale image-text pairs without relying on expensive manual annotations. However, existing methods often struggle with the noise inherent in web-collected data and the complexity of unstructured long medical texts. To address these challenges, we propose a novel VLP framework integrating a Multi-Agent data GENeration (MAGEN) system and Ontology-based Multi-Aspect Knowledge-Enhanced (O-MAKE) pretraining. First, MAGEN enhances data quality by synthesizing knowledge-enriched descriptions via a foundation model-assisted captioning and retrieval-based verification pipeline. Second, O-MAKE addresses the difficulty of learning from long, unstructured texts by decomposing them into distinct knowledge aspects. This facilitates fine-grained alignment at both global and patch levels, while explicitly modeling medical concept relationships through ontology-guided mechanisms. We validate our framework in the field of dermatology, where comprehensive experiments demonstrate the effectiveness of each component. Our approach achieves state-of-the-art zero-shot performance on disease classification and cross-modal retrieval tasks across eight datasets. Our code and the augmented dataset Derm1M-AgentAug, comprising over 400k skin-image-text pairs, will be released at https://github.com/SiyuanYan1/Derm1M.

医学多模态知识增强生成模型皮肤病识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。