构建垃圾识别新基准,测试视觉大模型在复杂场景下的表现
Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
- 设计真实垃圾场景数据集,包含形变物体与复杂环境
- 发现当前视觉大模型在杂乱环境中准确率显著下降
- 适合研究模型鲁棒性与真实世界应用的学者参考
近期大语言模型(LLMs)的发展推动了视觉大语言模型(VLLMs)的进步,使其能够执行多种视觉理解任务。尽管LLMs在标准自然图像上表现优异,但在存在复杂环境和形变物体的杂乱数据集上的能力尚未充分探索。本文提出一个专为真实场景垃圾分类设计的新数据集,特征包括复杂环境与形变物体。同时,我们提出了深入的评估方法,用于严格检验VLLMs的鲁棒性和准确性。该数据集与全面分析为VLLM在挑战性条件下的性能提供了重要见解。研究结果强调了提升VLLM鲁棒性的迫切需求,以改善其在复杂环境中的表现。数据集与实验代码将公开共享。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have paved the way for Vision Large Language Models (VLLMs) capable of performing a wide range of visual understanding tasks. While LLMs have demonstrated impressive performance on standard natural images, their capabilities have not been thoroughly explored in cluttered datasets where there is complex environment having deformed shaped objects. In this work, we introduce a novel dataset specifically designed for waste classification in real-world scenarios, characterized by complex environments and deformed shaped objects. Along with this dataset, we present an in-depth evaluation approach to rigorously assess the robustness and accuracy of VLLMs. The introduced dataset and comprehensive analysis provide valuable insights into the performance of VLLMs under challenging conditions. Our findings highlight the critical need for further advancements in VLLM's robustness to perform better in complex environments. The dataset and code for our experiments will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。