arXiv:2602.13693cs.CV2026-02

用轻量微调技术生成糖尿病神经病变的角膜图像,提升医疗影像数据质量。

A WDLoRA-Based Multimodal Generative Framework for Clinically Guided Corneal Confocal Microscopy Image Synthesis in Diabetic Neuropathy

  • 基于权重分解低秩适配,分离神经走向与组织对比度学习
  • 合成图像在结构相似度上达0.630,视觉质量领先于主流模型
  • 生成数据可有效训练诊断模型,准确率提升2.1%

角膜共聚焦显微镜(CCM)是评估糖尿病周围神经病变(DPN)小纤维损伤的敏感工具,但深度学习诊断模型的发展受限于标注数据稀少和角膜神经形态的细微差异。尽管人工智能基础生成模型在自然图像合成中表现优异,但在医学影像中因领域特定训练不足,常导致解剖真实性欠缺。为此,我们提出一种基于权重分解低秩适应(WDLoRA)的多模态生成框架,实现临床引导下的CCM图像合成。该方法通过解耦权重更新的幅值与方向,使模型独立学习神经拓扑(方向)和基质对比度(强度),结合神经分割掩码与疾病特异性临床提示,生成涵盖正常、早期及进展期DPN的解剖一致图像。三重评估表明,该框架在视觉保真度(FID: 5.18)和结构完整性(SSIM: 0.630)上均达到当前最优,显著优于GAN与标准扩散模型。关键的是,合成图像保留了金标准临床生物标志物,统计上等同于真实患者数据。用于训练下游诊断模型时,合成数据使诊断准确率提升2.1%,分割性能提升2.2%,验证其缓解医学AI数据瓶颈的潜力。

原文摘要 · Abstract (English)

Corneal Confocal Microscopy (CCM) is a sensitive tool for assessing small-fiber damage in Diabetic Peripheral Neuropathy (DPN), yet the development of robust, automated deep learning-based diagnostic models is limited by scarce labelled data and fine-grained variability in corneal nerve morphology. Although Artificial Intelligence (AI)-driven foundation generative models excel at natural image synthesis, they often struggle in medical imaging due to limited domain-specific training, compromising the anatomical fidelity required for clinical analysis. To overcome these limitations, we propose a Weight-Decomposed Low-Rank Adaptation (WDLoRA)-based multimodal generative framework for clinically guided CCM image synthesis. WDLoRA is a parameter-efficient fine-tuning (PEFT) mechanism that decouples magnitude and directional weight updates, enabling foundation generative models to independently learn the orientation (nerve topology) and intensity (stromal contrast) required for medical realism. By jointly conditioning on nerve segmentation masks and disease-specific clinical prompts, the model synthesises anatomically coherent images across the DPN spectrum (Control, T1NoDPN, T1DPN). A comprehensive three-pillar evaluation demonstrates that the proposed framework achieves state-of-the-art visual fidelity (Fréchet Inception Distance (FID): 5.18) and structural integrity (Structural Similarity Index Measure (SSIM): 0.630), significantly outperforming GAN and standard diffusion baselines. Crucially, the synthetic images preserve gold-standard clinical biomarkers and are statistically equivalent to real patient data. When used to train automated diagnostic models, the synthetic dataset improves downstream diagnostic accuracy by 2.1% and segmentation performance by 2.2%, validating the framework's potential to alleviate data bottlenecks in medical AI.

医学影像生成模型糖尿病神经病变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。