arXiv:2601.00905cs.CVcs.AI2026-01

用视觉语言模型判断垃圾能否回收及如何分类

Evaluating Contextual Intelligence in Recyclability: A Comprehensive Study of Image-Based Reasoning Systems

  • 用图像和多模型评估垃圾回收分类能力
  • 模型能识别混合材质与破损物品的回收方式
  • 适合环保科技与智能垃圾分类研究者

尽管高效回收的重要性广受认可,但公众准确判断物品可回收性及正确处理仍具挑战。本研究探索GPT-4o、GPT-4o-mini和Claude 3.5等前沿视觉语言模型在预测常见废弃物回收性方面的应用。基于定制图像数据集,评估模型将物品匹配至合适回收箱的能力,包括判断物品是否可物理放入。同时考察模型在三类复杂场景下的表现:(i)依据地区差异调整回收规则;(ii)识别污染或结构损坏影响;(iii)处理多材料复合物体。结果表明,相比以往版本,这些模型在上下文理解上取得显著进步,但仍存在不足。持续优化具备上下文感知能力的模型,对提升公众回收行为与推动环境可持续发展至关重要。

原文摘要 · Abstract (English)

While the importance of efficient recycling is widely acknowledged, accurately determining the recyclability of items and their proper disposal remains a complex task for the general public. In this study, we explore the application of cutting-edge vision-language models (GPT-4o, GPT-4o-mini, and Claude 3.5) for predicting the recyclability of commonly disposed items. Utilizing a curated dataset of images, we evaluated the models' ability to match objects to appropriate recycling bins, including assessing whether the items could physically fit into the available bins. Additionally, we investigated the models' performance across several challenging scenarios: (i) adjusting predictions based on location-specific recycling guidelines; (ii) accounting for contamination or structural damage; and (iii) handling objects composed of multiple materials. Our findings highlight the significant advancements in contextual understanding offered by these models compared to previous iterations, while also identifying areas where they still fall short. The continued refinement of context-aware models is crucial for enhancing public recycling practices and advancing environmental sustainability.

视觉语言模型垃圾分类环保科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。