用视觉语言模型自动检测自动驾驶数据中的标注错误,提升数据质量。
AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
- 利用视觉语言模型识别图像标注中的错误,实现自动化清洗。
- 在KITTI和nuImages数据集上验证,错误检测率显著提升。
- 适合需要高质量标注数据的自动驾驶研发团队使用。
自动驾驶系统训练依赖大量精确标注的数据集以实现稳健性能。人工标注存在缺陷,往往需多次迭代才能获得高质量数据,但手动审查大规模数据集耗时且成本高。本文提出AutoVDC(自动化视觉数据清洗)框架,探索利用视觉语言模型(VLMs)自动识别视觉数据集中的错误标注,帮助用户消除错误并提升数据质量。我们在KITTI和nuImages数据集上验证该方法,这两个数据集包含自动驾驶的目标检测基准。为评估AutoVDC效果,我们人为注入错误标注生成数据变体,并测试不同VLM的错误检测率。此外,还比较了不同VLM的表现,并研究微调对管道性能的影响。结果表明,该方法在错误检测与数据清洗实验中表现优异,具备显著提升大规模生产数据集可靠性与准确性的潜力。
原文摘要 · Abstract (English)
Training of autonomous driving systems requires extensive datasets with precise annotations to attain robust performance. Human annotations suffer from imperfections, and multiple iterations are often needed to produce high-quality datasets. However, manually reviewing large datasets is laborious and expensive. In this paper, we introduce AutoVDC (Automated Vision Data Cleaning) framework and investigate the utilization of Vision-Language Models (VLMs) to automatically identify erroneous annotations in vision datasets, thereby enabling users to eliminate these errors and enhance data quality. We validate our approach using the KITTI and nuImages datasets, which contain object detection benchmarks for autonomous driving. To test the effectiveness of AutoVDC, we create dataset variants with intentionally injected erroneous annotations and observe the error detection rate of our approach. Additionally, we compare the detection rates using different VLMs and explore the impact of VLM fine-tuning on our pipeline. The results demonstrate our method's high performance in error detection and data cleaning experiments, indicating its potential to significantly improve the reliability and accuracy of large-scale production datasets in autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。