arXiv:2509.11201cs.CV2025-09IJCV被引 4

用合成数据训练森林树体分割模型,大幅减少真实标注数据需求。

Scaling Up Forest Vision with Synthetic Data

  • 用游戏引擎与物理模拟生成大规模3D森林合成数据,替代昂贵真实采集。
  • 仅需0.1公顷真实数据微调,模型性能接近全量真实数据训练结果。
  • 适合需要少标注数据的生态遥感、林业监测研究者使用。

精准的树木分割是從森林激光掃描中提取單株樹木指標的關鍵步驟,對理解碳循環等生態系統功能至關重要。過去十年間,受人工智能發展推動,樹木分割算法快速進步。然而現有的公開三維森林數據集規模不足,難以訓練魯棒的分割系統。受自動駕駛領域合成數據成功的啟發,本文探討合成數據在樹木分割中的應用可能性。我們採用合成數據進行預訓練,僅需少量真實森林樣地標註即可微調。開發了新的合成數據生成流程,融合遊戲引擎與基於物理的激光雷達模擬技術,構建了一個前所未有的大規模、多樣化、已標註三維森林數據集。實驗表明,使用該合成數據預訓練後,僅需在單個小於0.1公頃的真实森林樣地微調,模型分割效果即能與全量真實數據訓練模型相媲美。同時,我們識別出合成數據成功應用的關鍵因素:物理真實性、數據多樣性與規模。相關代碼與數據集已公開於https://github.com/yihshe/CAMP3D.git。

原文摘要 · Abstract (English)

Accurate tree segmentation is a key step in extracting individual tree metrics from forest laser scans, and is essential to understanding ecosystem functions in carbon cycling and beyond. Over the past decade, tree segmentation algorithms have advanced rapidly due to developments in AI. However existing, public, 3D forest datasets are not large enough to build robust tree segmentation systems. Motivated by the success of synthetic data in other domains such as self-driving, we investigate whether similar approaches can help with tree segmentation. In place of expensive field data collection and annotation, we use synthetic data during pretraining, and then require only minimal, real forest plot annotation for fine-tuning. We have developed a new synthetic data generation pipeline to do this for forest vision tasks, integrating advances in game-engines with physics-based LiDAR simulation. As a result, we have produced a comprehensive, diverse, annotated 3D forest dataset on an unprecedented scale. Extensive experiments with a state-of-the-art tree segmentation algorithm and a popular real dataset show that our synthetic data can substantially reduce the need for labelled real data. After fine-tuning on just a single, real, forest plot of less than 0.1 hectare, the pretrained model achieves segmentations that are competitive with a model trained on the full scale real data. We have also identified critical factors for successful use of synthetic data: physics, diversity, and scale, paving the way for more robust 3D forest vision systems in the future. Our data generation pipeline and the resulting dataset are available at https://github.com/yihshe/CAMP3D.git.

森林视觉合成数据点云分割3D建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。