arXiv:2506.12323cs.CV2025-06NeurIPS被引 12

用AI专家反馈生成更真实的皮肤病图像,提升诊断模型效果。

Doctor Approved: Generating Medically Accurate Skin Disease Images through AI-Expert Feedback

  • 用AI模拟医生评审,指导扩散模型生成更准确的皮肤图像。
  • 合成图像临床质量显著提升,少样本下诊断准确率提高13.89%。
  • 减少人工标注负担,适合医疗数据稀缺场景使用。

医学数据匮乏严重限制了诊断机器学习模型的泛化能力,因小规模临床数据无法覆盖疾病全貌。为解决此问题,扩散模型(DMs)被视为生成合成图像和数据增强的有前途方向。然而,现有方法常生成医学不准确的图像,损害模型性能。领域专家知识对正确编码临床信息至关重要,尤其在数据稀少、质量重于数量的情况下。现有结合人工反馈的方法如强化学习(RL)和直接偏好优化(DPO),依赖强奖励函数或需耗时的人工评估。近期多模态大语言模型(MLLMs)展现出强大的视觉推理能力,使其成为理想的评估者候选。本文提出新框架MAGIC(通过AI-专家协作生成医学准确图像),用于生成可用于数据增强的临床准确皮肤病图像。该方法将专家定义的标准转化为可操作的反馈,显著提升合成图像的临床准确性,同时降低人工工作量。实验表明,该方法大幅提高合成图像的临床质量,其输出与皮肤科医生评估一致。此外,用这些合成图像扩充训练数据,在20类皮肤病分类任务中诊断准确率提升9.02%,在少样本设置下提升13.89%。

原文摘要 · Abstract (English)

Paucity of medical data severely limits the generalizability of diagnostic ML models, as the full spectrum of disease variability can not be represented by a small clinical dataset. To address this, diffusion models (DMs) have been considered as a promising avenue for synthetic image generation and augmentation. However, they frequently produce medically inaccurate images, deteriorating the model performance. Expert domain knowledge is critical for synthesizing images that correctly encode clinical information, especially when data is scarce and quality outweighs quantity. Existing approaches for incorporating human feedback, such as reinforcement learning (RL) and Direct Preference Optimization (DPO), rely on robust reward functions or demand labor-intensive expert evaluations. Recent progress in Multimodal Large Language Models (MLLMs) reveals their strong visual reasoning capabilities, making them adept candidates as evaluators. In this work, we propose a novel framework, coined MAGIC (Medically Accurate Generation of Images through AI-Expert Collaboration), that synthesizes clinically accurate skin disease images for data augmentation. Our method creatively translates expert-defined criteria into actionable feedback for image synthesis of DMs, significantly improving clinical accuracy while reducing the direct human workload. Experiments demonstrate that our method greatly improves the clinical quality of synthesized skin disease images, with outputs aligning with dermatologist assessments. Additionally, augmenting training data with these synthesized images improves diagnostic accuracy by +9.02% on a challenging 20-condition skin disease classification task, and by +13.89% in the few-shot setting.

皮肤病生成扩散模型AI医生数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。