用合成数据提升油菜分枝计数,最优比例与标签分布可降误差14.7%。
The Effects of Synthetic Data and Label Distribution on Canola Branch Counting

- 用L系统生成合成图像,调节真实与合成数据比例及标签分布。
- 最佳比例1:7,标签分布趋近真实时误差降至0.912,比纯真实数据低14.7%。
- 每类标签至少10张合成图即可,100张反会降低性能,适合资源有限者。
为自动化表型分析收集标注植物图像常耗时且昂贵。利用模拟生长发育的植物模型可生成无限带精确标签的合成图像。然而,先前研究指出,引入合成数据能否提升性能取决于合成与真实图像的比例以及合成数据的标签分布。为系统量化这两个因素,我们使用校准的L系统植物模型,在油菜分枝计数任务上训练ResNet-18模型,并独立调节每个因素。合成/真实比例在1:5至1:22之间普遍提升性能;最佳比例1:7使平均绝对差降低7.6%。对于标签分布,均匀分布表现差(绝对误差约1.70),向真实分布插值90%后误差降至0.927;而对真实标签分布进行高斯平滑可得最优结果(绝对误差0.912,较纯真实数据提升14.7%)。每类标签最少10张合成图像即有小幅收益,100张则过拟合并损害性能。
原文摘要 · Abstract (English)
Collecting annotated plant images for automated phenotyping is often slow and expensive. Plant models simulating growth and development can generate unlimited synthetic images with exact labels. However, previous work has established that whether incorporating synthetic data improves performance depends on the ratio of synthetic to real images and the label distribution of the synthetic dataset. To systematically quantify both factors, we train ResNet-18 models on a canola branch-counting task using a calibrated L-system plant model. We vary each factor independently. Synthetic-to-real ratios of 1:5 to 1:22 broadly improve performance; the best ratio (1:7) reduces mean absolute difference by 7.6% over real-only training. For label distribution, a uniform synthetic distribution is strongly suboptimal (abs. diff. of approximately 1.70); interpolating 90% toward the real distribution yields abs. diff. 0.927, whereas Gaussian smoothing of the real label distribution yields the best overall result (abs. diff. 0.912, a 14.7% improvement over real-only). A minimum of 10 synthetic images per label offers a simpler alternative with modest gains, while 100 per label over-corrects and hurts performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。