用视觉问答技术提升垃圾分类合规性,助力智慧环保
WasteAssistant: Regulation-Guided Visual Question Answering Framework for Intelligent Waste Segregation and Sustainable Managemen

- 融合视觉与语言模型,实现多模态智能判别
- 在13,500组问答数据上,BLEU达0.8291,优于传统方法
- 贴合印度固废管理法规,适合城市环卫系统落地
高效垃圾分类对可持续城市发展和环境治理至关重要。现有自动化系统受限于单模态视觉处理、上下文理解不足及监管对齐弱。为此,我们提出一种语言引导的视觉-AI框架,整合视觉语言模型与多模态大语言模型,实现联合视觉-语言推理。该框架遵循印度《2016年固体废物管理规则》。构建了包含13,500个问答对的新数据集WasteVQA,涵盖21类垃圾。实验表明,基于BLIP的模型在该数据集上达到BLEU 0.8291和BERTScore 0.9273,优于传统CNN方法。本工作提升了源头分类准确率,确保监管合规,支持市政与市民端的可扩展部署,推动多模态AI在可持续城市基础设施中的应用。源代码与数据集已开源:https://github.com/Khushkataruka/WasteAssistant
原文摘要 · Abstract (English)
Efficient waste segregation is critical for sustainable urban management and environmental governance. Existing automated systems are limited by single-modality visual processing, insufficient contextual understanding, and weak regulatory alignment. To address these issues, we propose a language-guided vision-AI framework that integrates vision-language models and multimodal large language models for joint visual-linguistic reasoning. This framework implements a visual question answering paradigm aligned with India's Solid Waste Management Rules 2016. We construct a new WasteVQA dataset with 13,500 question-answer pairs across 21 waste categories. Experiments show that the BLIP-based model achieves a BLEU score of 0.8291 and a BERTScore of 0.9273, outperforming traditional CNN-based methods. This work improves source-level segregation accuracy, ensures regulatory compliance, and supports scalable deployment for municipal and citizen-facing waste management, promoting multimodal AI in sustainable urban infrastructure. The source code and dataset are available at: https://github.com/Khushkataruka/WasteAssistant
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。