公开10类垃圾图像数据集,助力智能垃圾分类研究。
The Garbage Dataset (GD): A Multi-Class Image Benchmark for Automated Waste Segregation
- 构建包含12259张图像的多类别垃圾数据集,覆盖10类常见垃圾。
- EfficientNetV2S模型在该数据集上达到95.13%准确率和0.95 F1-score。
- 数据集揭示了类别不平衡、背景复杂等现实挑战,适合可持续计算研究。
本研究提出公开可用的垃圾数据集(GD),旨在推动机器学习与计算机视觉在自动垃圾分类中的应用。该数据集涵盖金属、玻璃、生物、纸张、电池、普通垃圾、纸板、鞋子、衣物和塑料共10类常见家庭垃圾,包含通过DWaste移动应用及精选网络来源收集的12,259张标注图像。数据采集采用校验和与异常值检测进行严格验证,利用PCA/t-SNE分析类别不平衡与视觉可分性,通过熵与显著性度量评估背景复杂度。使用EfficientNetV2M、EfficientNetV2S、MobileNet、ResNet50、ResNet101等先进深度学习模型进行性能与运行碳排放评估。实验表明,EfficientNetV2S在95.13%准确率和0.95 F1-score下表现最佳,碳成本适中。分析发现数据集存在类别不平衡、塑料/纸板/纸张等高异常值类别偏倚以及亮度变化等问题,需在实际部署中重点关注。结论认为GD为垃圾分类研究提供了有价值的现实基准,同时揭示了类别不平衡、背景复杂性及模型选择中的环境权衡等关键挑战。数据集已公开发布,以支持环境可持续性相关研究。
原文摘要 · Abstract (English)
This study introduces the Garbage Dataset (GD), a publicly available image dataset designed to advance automated waste segregation through machine learning and computer vision. It is a diverse dataset that covers 10 categories of common household waste: metal, glass, biological, paper, battery, trash, cardboard, shoes, clothes, and plastic. The dataset comprises 12,259 labeled images collected through multiple methods, including the DWaste mobile app and curated web sources. The methods included rigorous validation through checksums and outlier detection, analysis of class imbalance and visual separability through PCA/t-SNE, and assessment of background complexity using entropy and saliency measures. The dataset was benchmarked using state-of-the-art deep learning models (EfficientNetV2M, EfficientNetV2S, MobileNet, ResNet50, ResNet101) evaluated on performance metrics and operational carbon emissions. The results of the experiment indicate that EfficientNetV2S achieved the highest performance with a accuracy of 95.13% and an F1-score of 0.95 with moderate carbon cost. Analysis revealed inherent dataset characteristics including class imbalance, a skew toward high-outlier classes (plastic, cardboard, paper), and brightness variations that require consideration. The main conclusion is that GD provides a valuable real-world benchmark for waste classification research while highlighting important challenges such as class imbalance, background complexity, and environmental trade-offs in model selection that must be addressed for practical deployment. The dataset is publicly released to support further research in environmental sustainability applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。