arXiv:2511.09604eess.IV2025-11

用扩散模型生成光伏缺陷图像,解决数据少难题。

Bridging the Data Gap: Spatially Conditioned Diffusion Model for Anomaly Generation in Photovoltaic Electroluminescence Images

论文配图:Bridging the Data Gap: Spatially Conditioned Diffusion Model for Anomaly Generation in Photovoltaic Electroluminescence Images
图 1 · 摘自论文原文
  • 基于空间条件扩散模型,按掩码控制缺陷位置与类型
  • 生成图像FID仅4.10,真实度高;训练效果提升8.34点
  • 适合光伏检测、缺陷生成与数据增强研究者

可靠检测光伏组件缺陷对保障太阳能效率至关重要,但高质量、多样且均衡的标注数据集稀缺制约了计算机视觉模型的发展。本文提出PV-DDPM,一种空间条件化去噪扩散概率模型,可针对多晶硅、单晶硅、半切多晶硅及背接触指状互连(狗骨型互联)四种电池类型生成电致发光(EL)图像中的异常。通过二值掩码控制结构特征与缺陷位置,实现单缺陷与多缺陷场景的可控合成。据我们所知,这是首个联合建模多种光伏电池类型并支持多种异常类型同步生成的框架。同时构建了增强版数据集E-SCDD,包含1,000张像素级标注的EL图像,涵盖30个语义类别,以及1,768张未标注合成样本。定量评估显示,生成图像在所有类别上的弗雷切特初始距离(FID)为4.10,核初始距离(KID)为0.0023±0.0007。在E-SCDD上训练的视觉-语言异常检测模型AA-CLIP,相比SCDD数据集,像素级AUC与平均精度分别提升1.70和8.34个百分点。

原文摘要 · Abstract (English)

Reliable anomaly detection in photovoltaic (PV) modules is critical for maintaining solar energy efficiency. However, developing robust computer vision models for PV inspection is constrained by the scarcity of large-scale, diverse, and balanced datasets. This study introduces PV-DDPM, a spatially conditioned denoising diffusion probabilistic model that generates anomalous electroluminescence (EL) images across four PV cell types: multi-crystalline silicon (multi-c-Si), mono-crystalline silicon (mono-c-Si), half-cut multi-c-Si, and interdigitated back contact (IBC) with dogbone interconnect. PV-DDPM enables controlled synthesis of single-defect and multi-defect scenarios by conditioning on binary masks representing structural features and defect positions. To the best of our knowledge, this is the first framework that jointly models multiple PV cell types while supporting simultaneous generation of diverse anomaly types. We also introduce E-SCDD, an enhanced version of the SCDD dataset, comprising 1,000 pixel-wise annotated EL images spanning 30 semantic classes, and 1,768 unlabeled synthetic samples. Quantitative evaluation shows our generated images achieve a Fréchet Inception Distance (FID) of 4.10 and Kernel Inception Distance (KID) of 0.0023 $\pm$ 0.0007 across all categories. Training the vision--language anomaly detection model AA-CLIP on E-SCDD, compared to the SCDD dataset, improves pixel-level AUC and average precision by 1.70 and 8.34 points, respectively.

缺陷生成扩散模型光伏检测数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。