arXiv:2508.00639cs.CV2025-08中稿 · iMIMIC - Interpret…被引 1

用20个标注样本生成假数据,提升肺结节诊断解释性模型性能

Minimum Data, Maximum Impact: 20 annotated samples for explainable lung nodule classification

  • 用扩散模型生成带属性标注的肺结节图像,仅需20个真实标注样本
  • 合成数据使属性预测准确率提升13.4%,诊断准确率提升1.8%
  • 适合医疗AI可解释性研究者,解决小样本标注难题

提供人类可理解解释的分类模型能增强临床医生对医学影像诊断中AI的信任与可用性。研究重点是将放射科医生使用的病理相关视觉属性(如形状、纹理)与诊断结果联合建模,使AI决策过程更贴近临床思维。然而,这类模型推广受限于缺乏大规模带属性标注的医学图像数据集。为应对这一挑战,我们提出利用生成模型合成属性标注数据。通过在扩散模型中引入属性条件,仅使用来自LIDC-IDRI数据集的20个带属性标注的肺结节样本进行训练。将生成图像加入可解释模型的训练后,属性预测准确率提高13.4%,目标诊断准确率提升1.8%。本工作展示了合成数据在克服数据局限方面的潜力,提升了可解释模型在医学图像分析中的适用性。

原文摘要 · Abstract (English)

Classification models that provide human-interpretable explanations enhance clinicians' trust and usability in medical image diagnosis. One research focus is the integration and prediction of pathology-related visual attributes used by radiologists alongside the diagnosis, aligning AI decision-making with clinical reasoning. Radiologists use attributes like shape and texture as established diagnostic criteria and mirroring these in AI decision-making both enhances transparency and enables explicit validation of model outputs. However, the adoption of such models is limited by the scarcity of large-scale medical image datasets annotated with these attributes. To address this challenge, we propose synthesizing attribute-annotated data using a generative model. We enhance the Diffusion Model with attribute conditioning and train it using only 20 attribute-labeled lung nodule samples from the LIDC-IDRI dataset. Incorporating its generated images into the training of an explainable model boosts performance, increasing attribute prediction accuracy by 13.4% and target prediction accuracy by 1.8% compared to training with only the small real attribute-annotated dataset. This work highlights the potential of synthetic data to overcome dataset limitations, enhancing the applicability of explainable models in medical image analysis.

肺结节可解释AI生成模型小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。