arXiv:2607.19535cs.CV2026-07

用合成图像提升校园垃圾识别模型,结果却不尽如人意。

Synthetic and Derived Training Images for Campus Waste Detection: A Multi-Seed Evaluation with YOLOv8n

论文配图:Synthetic and Derived Training Images for Campus Waste Detection: A Multi-Seed Evaluation with YOLOv8n
图 1 · 摘自论文原文
  • 用合成与衍生图像增强真实数据训练YOLOv8n模型
  • 真实数据单独训练达[email protected] 0.691,合成数据反而降低性能
  • 实验揭示种子差异影响显著,小样本难得出可靠结论

错误丢弃会污染校园回收流,摄像头安装在垃圾桶上可实时反馈投放行为。本文评估了合成与衍生图像是否能提升针对该视角的YOLOv8n检测器性能。真实数据集包含148张校园照片:86张用于训练,31张用于验证,31张用于测试。十二种联合训练配置改变了添加图像的数量与来源。对七种主要设置重复四次匹配种子,计算置信区间。仅使用真实数据的模型平均[email protected]为0.691 [0.665, 0.722];背景替换后降至0.560 [0.499, 0.619],孤立物体图像得0.680 [0.644, 0.724],完整增强池则为0.487 [0.438, 0.537]。还测试了手部与前臂组合图像,因初始组合含测试图,故重做并重新运行四次种子。修正后的配对差异为+0.034 [-0.063, 0.199],不支持手部复合有显著效果。单种子迁移实验显示混合与分步预训练的排名依赖源数据。所有配置均未超越仅真实数据基线。报告区间量化了种子变异;31张测试集仍不足以得出强分类结论。

原文摘要 · Abstract (English)

Incorrect disposal can contaminate campus recycling streams, and a bin-mounted camera could provide feedback as an item is discarded. We evaluated whether synthetic and derived images improve a YOLOv8n detector for this view. The real dataset contained 148 campus photographs: 86 for training, 31 for validation, and 31 for testing. Twelve joint-training configurations varied the amount and source of added images. We repeated seven principal settings with four matched seeds and computed bootstrap percentile intervals over those seeds. The real-only model reached a mean [email protected] of 0.691 [0.665, 0.722]. Background replacement reduced the mean to 0.560 [0.499, 0.619], isolated-object images gave 0.680 [0.644, 0.724], and the full augmentation pool gave 0.487 [0.438, 0.537]. We also tested hand-and-forearm composites because every real photo showed a held object. Two cutouts in the initial composite set came from test photographs, so we discarded that experiment, rebuilt the set with training-split cutouts, and reran all four seeds. The corrected paired difference was +0.034 [-0.063, 0.199], which does not support a reliable hand-composite effect. Single-seed transfer experiments produced source-dependent rankings between joint mixing and sequential pretraining. None of the evaluated configurations exceeded the real-only baseline. The reported intervals quantify seed variation; the 31-photo test set remains too small for strong class-specific conclusions.

目标检测合成数据YOLOv8n校园垃圾

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。