arXiv:2605.05627cs.CVcs.AI2026-05被引 1

用AI生成图像解决森林再生数据少难题,提升细粒度分割精度。

Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping

论文配图:Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping
图 1 · 摘自论文原文
  • 用大模型生成高保真图像与语义掩码,替代部分人工标注。
  • 融合真实与生成数据训练,整体F1提升超15个百分点。
  • 少量生成数据即可显著改善稀有类别的识别效果,适合小样本场景。

可持续森林管理依赖精准的物种组成制图,但传统地面调查耗时且地理覆盖有限。尽管无人机可规模化采集数据,深度学习解析仍受限于专家标注图像严重不足,尤其在视觉异质性强的再生区域。本文提出一种可扩展框架,减少对人工影像解译的依赖,针对高分辨率、毫米级航拍图像进行细粒度语义分割。关键在于利用大规模Nano Banana Pro模型,从提示词同步生成高保真图像及其像素对齐的语义掩码。我们引入WilDReF-Q-V2,扩充自然林数据集,包含13,977张未标注和50张手工标注的真实图像,并构建了包含2101对合成图像与语义掩码的Gen4Regen数据集。方法融合真实与生成数据,表明生成数据与真实数据高度互补,联合训练使F1分数相比纯监督基线提升超过15个百分点。此外,即使少量提示生成数据也能显著提升少数类别性能,部分类别每类F1提升超30个百分点。结论表明,大规模视觉模型可作为敏捷数据生成器,有效启动稀缺专家标签领域的感知任务。数据集、代码与模型将公开于https://norlab-ulaval.github.io/gen4regen。

原文摘要 · Abstract (English)

Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained. While Uncrewed Aerial Vehicles (UAVs) offer scalable data collection, the transition to deep learning-based interpretation is bottlenecked by the severe scarcity of expert-annotated imagery, particularly in complex, visually heterogeneous regeneration zones. This paper addresses the dual challenges of data scarcity and extreme class imbalance in the fine-grained semantic segmentation of plants by providing a scalable framework that reduces reliance on manual photo-interpretation for high-resolution, millimetre-level aerial imagery. Importantly, we leverage the large-scale Nano Banana Pro model to simultaneously generate high-fidelity images and their corresponding pixel-aligned semantic masks from prompts. We introduce WilDReF-Q-V2, an expansion of a natural forest dataset with 13 977 new unlabelled and 50 hand-labelled real images, as well as the Gen4Regen dataset, featuring 2101 pairs of synthetic images and semantic masks. Our methodology integrates real-world data with AI-generated images, highlighting that AI-generated data is highly complementary to real-world data, with unified training yielding an F1 score improvement of over 15 %pt compared to purely supervised baselines. Furthermore, we demonstrate that even small quantities of prompt-generated data significantly improve performance for underrepresented classes, some of which see per-class F1 score gains of over 30 %pt. We conclude that large-scale vision models can serve as agile data generators, effectively bootstrapping perception tasks for niche AI domains where expert labels are scarce or unavailable. Our datasets, source code, and models will be available at https://norlab-ulaval.github.io/gen4regen.

图像生成语义分割数据增强林业遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。