用扩散模型生成稀有堤坝沙沸缺陷图像,解决标注数据少的问题。
Multi-Conditioned Diffusion Synthesis of Sand Boils for Low-Resource Earthen-Levee Inspection

- 基于多分支控制网和微调的稳定扩散,从少量真实样本生成合成图像。
- 生成1020张合成图,815张通过质量过滤,保持缺陷真实性和场景一致性。
- 提供可复现的生成方案,适合低资源场景下的缺陷检测数据增强。
土质堤坝上的沙沸是危及安全的缺陷,但像素级检测受限于标注数据稀缺。本文提出一种基于扩散模型的合成流水线,用于低资源沙沸图像生成。使用微调后的Stable Diffusion XL与多分支ControlNet堆叠,结合小规模精选参考集生成合成检查图像。采用软掩码修复协议,在保留真实缺陷像素的同时重绘周围场景,避免拼接痕迹与色彩偏移。掩码条件控制网可在指定掩码内生成新沙沸,使掩码自然成为分割标签;但由于大规模标签认证仍不可行,故默认发布软掩码预设。文本条件由基于分类体系的Prompt Atlas提供,可无代码扩展至新缺陷类别。从真实训练图像中生成1020张合成候选图,其中815张通过CLIP可接受性筛选。评估采用分布性、保真度-多样性指标,对比真实参考集与泊松基线,并审计分布外漂移与记忆现象。无单一预设占优,各预设在保真度、多样性与标签可靠性间权衡。因此,默认发布标签可靠的预设,将精心混合的集合视为自然增强集。本工作仅限图像质量、标签来源与多样性验证,下游分割留待未来。代码与人工制品清单已公开以保障可复现性。
原文摘要 · Abstract (English)
Sand boils on earthen levees are safety-critical defects, but pixel-level detection is limited by scarce annotations. We present a diffusion-based synthesis pipeline for low-resource sand-boil imagery. Using Stable Diffusion XL fine-tuned with DreamBooth and conditioned by a multi-branch ControlNet stack, the pipeline generates synthetic inspection images from a small curated reference set. A soft-mask inpainting protocol preserves the real defect pixels while re-rendering the surrounding scene, avoiding seams and color shifts from prior seamless-cloning compositing. A mask-conditioned ControlNet can also generate a new boil inside a chosen mask, making the mask the segmentation label by construction; however, because large-scale label certification remains unresolved with the available real-trained gate, we release the soft-mask preset as the default. Text conditioning is supplied by a taxonomy-driven Prompt Atlas that expands one domain specification into a stratified, CLIP-validated prompt bank and transfers to new defect classes without code changes. From the real training images, the pipeline produces 1,020 synthetic candidates, of which 815 pass a CLIP admissibility filter. We evaluate image quality using distributional and fidelity-diversity measures against the real reference set and a Poisson baseline, and audit for out-of-distribution drift and memorization. No single preset dominates; each trades off fidelity, diversity, and label reliability. We therefore release the label-reliable preset as the default and treat a curated mixture as the natural augmentation set. Our claims are limited to image quality, label provenance, and diversity; downstream segmentation is left for future work. Code and an artifact manifest are released for reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。