用3D渲染+扩散模型生成逼真蘑菇图像,解决农业视觉数据难标注问题。
A Scalable Pipeline Combining Procedural 3D Graphics and Guided Diffusion for Photorealistic Synthetic Training Data Generation in White Button Mushroom Segmentation
- 结合Blender 3D渲染与约束扩散模型,自动生成带标注的逼真图像。
- 在6000张合成图像上训练的模型零样本测试达F1=0.859,超越现有方法。
- 无需专业图形技能,可扩展至其他农作物检测场景。
工业蘑菇种植依赖计算机视觉进行监控与自动化采收,但准确的检测与分割模型需大量精确标注数据,成本高昂。合成数据提供可扩展替代方案,但常缺乏真实感。本文提出一种新工作流:将Blender中的3D渲染与约束扩散模型结合,自动生成高质、带标注的白帽菇(Agaricus Bisporus)照片级真实合成图像。该方法保持对3D场景配置和标注的完全控制,且无需特殊图形设计技能即可实现逼真效果。我们发布两个合成数据集(各含6000张图像,涵盖超25万株蘑菇实例),并评估基于这些数据训练的Mask R-CNN模型在零样本设置下的表现。在两个独立真实数据集(包括新采集基准)上测试,模型取得当前最优分割性能(在M18K上F1=0.859),仅使用合成数据即达成。尽管以白帽菇为例,该流程可轻松适配其他蘑菇种类或农业领域,如果实与叶片检测。
原文摘要 · Abstract (English)
Industrial mushroom cultivation increasingly relies on computer vision for monitoring and automated harvesting. However, developing accurate detection and segmentation models requires large, precisely annotated datasets that are costly to produce. Synthetic data provides a scalable alternative, yet often lacks sufficient realism to generalize to real-world scenarios. This paper presents a novel workflow that integrates 3D rendering in Blender with a constrained diffusion model to automatically generate high-quality annotated, photorealistic synthetic images of Agaricus Bisporus mushrooms. This approach preserves full control over 3D scene configuration and annotations while achieving photorealism without the need for specialized computer graphics expertise. We release two synthetic datasets (each containing 6,000 images depicting over 250k mushroom instances) and evaluate Mask R-CNN models trained on them in a zero-shot setting. When tested on two independent real-world datasets (including a newly collected benchmark), our method achieves state-of-the-art segmentation performance (F1 = 0.859 on M18K), despite using only synthetic training data. Although the approach is demonstrated on Agaricus Bisporus mushrooms, the proposed pipeline can be readily adapted to other mushroom species or to other agricultural domains, such as fruit and leaf detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。